Equipment control method and device, voice processing equipment, storage medium and computer program product
By combining action collection information and device processing results in smart home devices for dynamic control, the problem of difficult balance between power consumption and response speed of traditional voice control devices is solved, and more efficient energy management and rapid response are achieved.
Patent Information
- Application Number
- CN202510262040.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
In the field of smart homes, traditional voice control devices are difficult to balance between power consumption and response speed, resulting in high energy consumption or slow response speed.
By acquiring the action acquisition information of the device, it is determined whether it exceeds the preset threshold. If it exceeds the action, the action recognition process is performed. If the recognition result matches the target action, the networking process is performed, otherwise the device processing result is updated. This method combines action acquisition information and equipment processing results for dynamic control, giving priority to reducing unnecessary networking processing and energy consumption.
The power consumption and response speed of the device are balanced, the timeliness of the device are improved, and the power consumption of the device is reduced.
Smart Images

Figure CN120220671A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of smart home, and particularly to a device control method, device, voice processing device, storage medium and computer program product. Background Art
[0002] In the field of smart home, voice control functions have been widely applied. For example, users can turn on lights, adjust light brightness, start a home theater, control air conditioners, switch audio-visual channels, etc. through voice, which has taken a big step forward in the function of home automation.
[0003] In the traditional way, through the voice recognition function, the voice processing device can understand human speech and complete corresponding operations according to human oral commands, realizing direct language communication and control between humans and machines. For example, after waking up the smart speaker with a voice wake-up word, the voice signal is collected. However, this method requires the smart speaker to be powered on for a long time so that the smart speaker can monitor the user's voice in real time. If you want to reduce power consumption, the smart device needs to remain in the sleep mode when idle, and only provide relevant functions to the user after the user actively triggers it. However, this method also greatly reduces the response speed to user commands.
[0004] Therefore, how to achieve a balance between power consumption and response speed has become one of the problems to be solved urgently at present. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, voice processing device, computer-readable storage medium and computer program product that can achieve a balance between power consumption and response speed.
[0006] In a first aspect, the present application provides a device control method. The method includes:
[0007] If the action acquisition information of the device exceeds a first preset threshold, obtain the device processing result; the action acquisition information is used to represent the sensor data measured when the device is triggered; the device processing result is used to represent the historical failure result of the device executing the target processing; the target processing includes at least one of action recognition processing and voice processing;
[0008] If the device processing result exceeds a second preset threshold, perform the action recognition processing based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform network connection processing;
[0009] If the device processing result does not exceed the second preset threshold, perform the network connection processing to obtain a network connection result.
[0010] In one embodiment, after performing the action recognition process based on the information collected from the action and obtaining the action recognition result, the following steps are further included:
[0011] If the recognized action represented by the action recognition result matches the target action, then perform the networking process to obtain the networking result;
[0012] If the recognized action represented by the action recognition result does not match the target action, then update the device processing result.
[0013] In one embodiment, performing the networking process to obtain the networking result includes:
[0014] Collect the first voice signal and perform a networking process with the target device to obtain the networking result with the target device.
[0015] In one embodiment, after performing the networking process to obtain the networking result, the following steps are further included:
[0016] If the networking result indicates successful networking, then send the first voice signal to the target device so that the target device can recognize the first voice signal;
[0017] If the networking result indicates failed networking, then update the device processing result.
[0018] In one embodiment, after sending the first voice signal to the target device, the following steps are further included:
[0019] Set a first timer;
[0020] Within the first set time corresponding to the first timer, if the signal recognition result returned by the target device for the first voice signal is obtained and the signal recognition result indicates that the first voice signal is valid, then initialize the device processing result.
[0021] In one embodiment, the first timer is further used to indicate that within the first set time, perform an action recognition process on the currently acquired action information to obtain the current action recognition result;
[0022] After setting the first timer, the following steps are further included:
[0023] Within the first set time of the first timer, if it is detected that the current action represented by the current action recognition result matches the target action, then collect a second voice signal;
[0024] Send the second voice signal to the target device so that the target device can recognize the second voice signal;
[0025] Obtain the signal recognition result for the second voice signal returned by the target device.
[0026] In one embodiment, after setting the first timer, it further includes:
[0027] When the first set time of the first timer is reached, if the current action represented by the current action recognition result does not match the target action, and the signal recognition result indicating the validity of the first voice signal is not obtained, update the device processing result and end the networking process.
[0028] In one embodiment, the method further includes:
[0029] Set a second timer;
[0030] When the second set time corresponding to the second timer is reached, obtain the interrupt generated by the second timer; the interrupt is used to indicate initializing the device processing result.
[0031] In a second aspect, the present application also provides a device control device. The device includes:
[0032] An action measurement module, configured to obtain the current device processing result if the action acquisition information of the device exceeds a first preset threshold; the action acquisition information is used to represent the sensor data measured when the device is triggered; the device processing result is used to represent the historical failure result of the device performing target processing; the target processing includes at least one of action recognition processing and voice processing;
[0033] An action recognition module, configured to perform the action recognition processing based on the action acquisition information to obtain an action recognition result if the device processing result exceeds a second preset threshold; the action recognition result is used to indicate whether to perform the networking process;
[0034] A networking processing module, configured to perform the networking process to obtain a networking result if the device processing result does not exceed the second preset threshold.
[0035] In a third aspect, the present application also provides a voice processing device. The device includes a sensor and a controller;
[0036] The sensor is used to obtain the motion acquisition information of the voice processing device. If the motion acquisition information exceeds a first preset threshold, it instructs the controller to obtain the current device processing result; the motion acquisition information is used to represent the sensor data measured when the voice processing device is triggered; the device processing result is used to represent the historical failure result of the voice processing device performing target processing; the target processing includes at least one of motion recognition processing and voice processing.
[0037] The controller is used to perform the motion recognition processing on the received motion acquisition information to obtain a motion recognition result if it detects that the device processing result exceeds a second preset threshold; the motion recognition result is used to indicate whether to perform network connection processing; if the device processing result does not exceed the second preset threshold, it performs the network connection processing to obtain a network connection result.
[0038] In one embodiment, the controller is further used to control the sensor to enter a first working mode and the controller to enter a second working mode after the voice processing device is started; the power consumption of the second working mode is lower than that of the first working mode.
[0039] The sensor is further used to obtain the motion acquisition information of the voice processing device in the first working mode. When the motion acquisition information exceeds the first preset threshold, it wakes up the controller from the second working mode and instructs the controller to obtain the device processing result.
[0040] In one embodiment, the voice processing device further includes a wireless communication unit.
[0041] The controller is further used to control the wireless communication unit to perform the network connection processing to obtain the network connection result if the recognition action represented by the motion recognition result matches the target action; if the recognition action represented by the motion recognition result does not match the target action, it increments the device processing result by one and enters the second working mode again according to the set timing information.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a controller, the following steps are implemented:
[0043] If it is obtained that the motion acquisition information of the device exceeds a first preset threshold, obtain the device processing result; the motion acquisition information is used to represent the sensor data measured when the device is triggered; the device processing result is used to represent the historical failure result of the device performing target processing; the target processing includes at least one of motion recognition processing and voice processing.
[0044] If the processing result of the device exceeds a second preset threshold, perform the action recognition process based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform network connection processing;
[0045] If the processing result of the device does not exceed the second preset threshold, perform the network connection processing to obtain a network connection result.
[0046] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a controller, the following steps are implemented:
[0047] If the action acquisition information of the device obtained exceeds a first preset threshold, obtain the device processing result; the action acquisition information is used to characterize the sensor data measured when the device is triggered; the device processing result is used to characterize the historical failure result of the device performing a target process; the target process includes at least one of action recognition processing and voice processing;
[0048] If the processing result of the device exceeds a second preset threshold, perform the action recognition process based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform network connection processing;
[0049] If the processing result of the device does not exceed the second preset threshold, perform the network connection processing to obtain a network connection result.
[0050] The above device control method, device, voice processing device, storage medium, and computer program product, if the action acquisition information of the device exceeds the first preset threshold, then obtain the device processing result; the action acquisition information is used to represent the sensor data measured when the device is triggered; the device processing result is used to represent the historical failure result of the device executing the target processing; the target processing includes at least one of action recognition processing and voice processing; if the device processing result exceeds the second preset threshold, then perform action recognition processing based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform network connection processing; if the device processing result does not exceed the second preset threshold, then perform network connection processing to obtain a network connection result. By using this method, through combining the action acquisition information of the device and the device processing result for judgment and decision-making, if the device processing result exceeds the threshold, it indicates that there is a high failure rate or mis-trigger rate when the device executes the target processing. At this time, action recognition processing is preferentially performed based on the action acquisition information to reduce unnecessary network connection processing and energy consumption. If the device processing result does not exceed the threshold, then directly perform network connection processing to provide a quick response. Through this dynamic control mechanism, both the response timeliness of the device is improved and the power consumption of the device is reduced, thus effectively balancing the power consumption of the device and the response speed of network connection processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is an application environment diagram of the device control method in an embodiment;
[0052] Figure 2 It is a flowchart of the device control method in an embodiment;
[0053] Figure 3 It is a flowchart of the steps after obtaining the network connection result in an embodiment;
[0054] Figure 4 It is a flowchart of the device control method in another embodiment;
[0055] Figure 5 It is a structural block diagram of the device control device in an embodiment;
[0056] Figure 6 It is an internal structure diagram of the voice processing device in an embodiment;
[0057] Figure 7 It is a schematic diagram of the hardware system of the voice processing device in an embodiment;
[0058] Figure 8 It is a flowchart of the device control method in yet another embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0061] The device control method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the application environment can be a smart home system, and the smart home system can include: a router, a gateway, a terminal 101, a server 102, a device 103, etc. In the smart home system, the device 103 can be connected to the gateway through ZIGBEE / Bluetooth / WiFi, and the gateway and the user terminal can be respectively connected to the router through WiFi. In addition, the terminal can also establish a network connection with the server through 2G / 3G / 4G / 5G, WiFi, etc., so as to obtain the data sent by the server.
[0062] Among them, the terminal is used to provide local services for users; the terminal includes but is not limited to various mobile phones, tablet computers, laptop computers, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers, or can also be implemented by a cloud server. The device includes but is not limited to various voice control devices; for example, the device can be a small-sized Bluetooth voice remote control, or a voice companion.
[0063] In one embodiment, as Figure 2 shown, a device control method is provided. Taking the method applied to the Figure 1 device as an example for description. In this embodiment, the method includes the following steps:
[0064] Step S201, if the action acquisition information of the device exceeds the first preset threshold, then obtain the device processing result; the action acquisition information is used to represent the sensor data measured when the device is triggered; the device processing result is used to represent the historical failure result of the device executing the target processing; the target processing includes at least one of action recognition processing and voice processing.
[0065] Among them, the action acquisition information is used to describe the information about the action triggered for the device. For example, the action acquisition information can be the acceleration information or the angle information of the device when the user shakes (wiggles) the device. In practical applications, the action acquisition information can be the ACC value collected by the 3-axis accelerometer (Accelerometer, ACC) inside the 6-axis inertial sensor (6 IMU) in the device (such as a voice companion) and / or the angle information collected by the 3-axis gyroscope sensor, where the ACC value refers to the acceleration information of the device in the X, Y, and Z axis directions.
[0066] Among them, the first preset threshold is used to determine whether the device is triggered. In practical applications, the first preset threshold can be the interruption threshold set for the 6 IMU.
[0067] Among them, the historical failure result is used to describe the cumulative failure situation of the device in performing the target processing within the historical time period. For example, the historical failure result can include the historical failure times of the device in performing the target processing.
[0068] Specifically, configure parameters such as the first preset threshold for the device. When the device is not triggered, make the device work in the low-power mode and make the device enter the sleep state to save energy. It should be noted that even when the device is in the sleep state, the sensors in the device (such as the 6 IMU) will continuously detect the action acquisition information of the device. If the device detects that the action acquisition information exceeds the first preset threshold, it means that the action when the device is triggered is relatively large. At this time, the device (or a part of the components in the device, such as waking up the controller in the device) can be woken up, and then the device (or the controller) obtains the device processing result of the device in performing the target processing within the historical time period. For example, the device processing result can be the cumulative failure times of the device in performing action recognition processing and voice processing within the historical time period. If the device detects that the action acquisition information does not exceed the first preset threshold, there is no need to perform the subsequent step S202, and the device continues to maintain the sleep state to save energy.
[0069] Step S202, if the device processing result exceeds the second preset threshold, perform action recognition processing based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform network connection processing.
[0070] Among them, the second preset threshold is the judgment threshold set for the historical failure result of the device.
[0071] Among them, the action recognition result is used to indicate whether the action triggered for the device this time matches the target action.
[0072] Specifically, the device compares the device processing result with a second preset threshold. If it is detected that the device processing result exceeds the second preset threshold (for example, the number of historical failures Fail in the device processing result > the second preset threshold of 5 times), it indicates that the cumulative number of failures in the device's execution of the target processing during the historical time period is relatively large, and the credibility of the action triggered for the device this time is relatively low. It may be an accidental mis-triggering of the device by the user. To reduce the energy consumption caused by direct networking, the device first performs action recognition processing on the action acquisition information to identify what specific action was performed on the device when the device was triggered this time. Then, the device obtains the recognized action, and the device can also set the recognized action as the action recognition result. After that, the device can further determine whether to perform networking processing or other processing based on the action recognition result.
[0073] Step S203, if the device processing result does not exceed the second preset threshold, then perform networking processing to obtain a networking result.
[0074] Among them, the networking result is used to reflect the communication connection result between the device and the target device (such as a gateway). For example, the networking result can be a successful network connection or a failed network connection.
[0075] Specifically, if it is detected that the device processing result does not exceed the second preset threshold, it indicates that the cumulative number of failures in the device's execution of the target processing during the historical time period is relatively small, and the credibility of the action triggered for the device this time is relatively high. Then, the device can directly initiate networking processing to establish a communication connection with the target device, and then the device receives the networking result feedback by the target device.
[0076] In the above device control method, if the action acquisition information of the device exceeds the first preset threshold, then obtain the device processing result; the action acquisition information is used to characterize the sensor data measured when the device is triggered; the device processing result is used to characterize the historical failure result of the device in executing the target processing; the target processing includes at least one of action recognition processing and voice processing; if the device processing result exceeds the second preset threshold, then perform action recognition processing based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform networking processing; if the device processing result does not exceed the second preset threshold, then perform networking processing to obtain a networking result. By using this method, through the combination of the action acquisition information of the device and the device processing result for judgment and decision-making, if the device processing result exceeds the threshold, it indicates that there is a relatively high failure rate or mis-triggering rate when the device executes the target processing. At this time, priority is given to performing action recognition processing based on the action acquisition information to reduce unnecessary networking processing and energy consumption. If the device processing result does not exceed the threshold, then directly perform networking processing to provide a quick response. Through this dynamic control mechanism, both the response timeliness of the device is improved and the power consumption of the device is reduced, thus effectively balancing the power consumption of the device and the response speed of the networking processing.
[0077] In one embodiment, after performing action recognition processing based on the action acquisition information to obtain an action recognition result in step S202 above, the method further includes: if the recognized action represented by the action recognition result matches the target action, perform networking processing to obtain a networking result; if the recognized action represented by the action recognition result does not match the target action, update the device processing result.
[0078] The target action refers to a specific action set that can trigger the device to perform networking processing. For example, the target action can be "picking up" the device or "shaking" the device.
[0079] Specifically, if the recognized action identified based on the action acquisition information is the same as any one of the set target actions (such as "picking up" or "shaking"), the device can confirm that the recognized action triggered this time conforms to the specific target action, and then the device can directly perform networking processing, establish a communication connection with the target device, and then the device receives the networking result feedback by the target device.
[0080] Further, if the recognized action identified based on the action acquisition information is different from all the target actions, the device can confirm that the recognized action triggered this time does not conform to the specific target action, and determine that the device is triggered by mistake. Then the device updates the device processing result, increases the historical failure result represented by the device processing result. For example, increment the historical failure count (Fail) in the historical failure result by one; initialize the historical success count. The device can also set a second timer (such as a sleep timer) and a second setting time for the second timer. When the second setting time is reached, the device shuts down the second timer and initializes both the historical failure count and the historical success count in the historical failure result represented by the device processing result, and the device returns to the sleep state.
[0081] For example, if the recognized action does not match the target action, the device can set a 5-minute sleep timer, increment the historical failure count Fail by 1, and initialize the historical success count Successful to 0. After 5 minutes, the device initializes both Successful and Fail to 0, and then the device enters the sleep state.
[0082] In this embodiment, by determining whether the recognized action matches the target action that triggers networking processing, it is determined whether to perform networking processing or update the failure count in the device processing result, accurately identifying the interference of mis-triggered actions, thereby effectively reducing unnecessary networking processing and energy consumption. While achieving the balance between power consumption and response speed, it also improves the accuracy of device response.
[0083] In one embodiment, in step S203 above, network connection processing is performed to obtain a network connection result, which specifically includes the following: collecting a first voice signal and performing network connection processing with a target device to obtain a network connection result with the target device.
[0084] Among them, the target device refers to other devices with which the device wants to establish a communication connection. For example, the target device can be a gateway.
[0085] Among them, the first voice signal refers to the voice signal collected by the device during the network connection process. The voice signal can carry the voice command information of the user. For example, the voice signal can be "turn on the socket".
[0086] Specifically, when the device performs network connection processing with the target device, it can first activate the Bluetooth of the device and connect to the target device through Bluetooth; then the device receives the network connection result feedback by the target device. During the process of performing network connection processing, the device can also synchronously collect voice signals. To distinguish it from the voice signals collected in the subsequent steps, the voice signals collected here can be marked as the first voice signals.
[0087] In this embodiment, the device performs network connection processing with the target device and synchronously collects the first voice signal, so that the subsequent steps can transmit the collected first voice signal immediately after the network connection is successful, effectively saving the time for separately collecting voice signals, thereby improving the response speed of the device.
[0088] To further save energy consumption, when the device performs network connection processing with the target device, the interrupt of the acceleration sensor and the power supply of the IMU can also be turned off and turned on when needed.
[0089] In one embodiment, as Figure 3 shown, after performing network connection processing in step S203 above to obtain a network connection result, the following is further included:
[0090] Step S301, if the network connection result indicates a successful network connection, then send the first voice signal to the target device so that the target device can recognize the first voice signal.
[0091] Specifically, the device receives the network connection result returned by the target device; if it is detected that the network connection result indicates a successful network connection, the device can send the first voice signal collected during the network connection process to the target device through the communication connection with the target device for the target device to analyze the semantics represented by the first voice signal and perform corresponding actions.
[0092] For example, assume that the first voice signal collected by the device is "Turn on the living room light". After successful networking, the device sends the first voice signal "Turn on the living room light" to the target device through the communication connection. After receiving the first voice signal, the target device recognizes the first voice signal and performs the action of turning on the living room light in the voice recognition result.
[0093] Furthermore, the device can also set the maximum voice duration of the first voice signal that can be sent. For example, the maximum voice duration of the first voice signal can be set to 10s. Then, the first voice signal within 10s and 10s can be fully sent to the target device. If the first voice signal exceeds 10s, only the first 10s of the first voice signal will be sent, and the part of the first voice signal that exceeds 10s will automatically stop being sent and be considered sent successfully.
[0094] Step S302, if the networking result indicates a networking failure, update the device processing result.
[0095] Specifically, if it is detected that the networking result indicates a networking failure, which means that the number of failed attempts for the device to perform the target processing increases again, the device updates the device processing result in real time, increasing the historical failure result represented by the device processing result. For example, increment the historical failure count (Fail) in the historical failure result by one. The device can also set a sleep timer, and when the time corresponding to the sleep timer is reached, the device returns to the sleep state.
[0096] For example, if the networking fails, the device can increment the historical failure count Fail by 1 and initialize the historical success count Successful to 0.
[0097] In this embodiment, when the networking is successful, the first voice signal is transmitted to the target device to enable the target device to recognize the first voice signal and execute the content of the first voice signal. When the networking fails, the device processing result is updated in real time, providing a reliable basis for subsequent device trigger judgment decisions. In addition, the power consumption of the device for transmitting the voice signal can be controlled by setting the maximum voice duration of the first voice signal, avoiding excessive energy consumption of the device due to an overly long first voice signal, effectively reducing the transmission power consumption of the device, and extending the service life of the device.
[0098] In one embodiment, after the above step S301 of sending the first voice signal to the target device, it further includes: setting a first timer; within the first set time corresponding to the first timer, if the signal recognition result for the first voice signal returned by the target device is obtained and the signal recognition result indicates that the first voice signal is valid, initialize the device processing result.
[0099] Among them, the first setting time is used to represent the duration for which the device waits for the target device to feedback the signal recognition result. For example, the first setting time can be set to 3s, and the device will continuously wait for the target device to feedback the signal recognition result within 3s.
[0100] Among them, the voice signal is valid. "Valid" means that the semantic meaning of the voice signal can be recognized.
[0101] Specifically, after the device finishes sending the first voice signal, it can also start a timer and set the timing duration, and then the device obtains the first timer and the first setting time corresponding to the first timer. Within the first setting time, the device will wait for the target device to feedback the signal recognition result for the first voice signal. If the device receives the signal recognition result returned by the target device for the first voice signal within the first setting time, and the signal recognition result indicates that the first voice signal is a valid voice command, then the device can initialize the device processing result. For example, initialize the historical failure count (Fail) in the device processing result to 0, and increase the historical success count. For example, increase the historical success count (Successful) by 1; otherwise, the device will continue to wait for the target device to feedback the signal recognition result for the first voice signal until the first setting time is reached or a target action triggered for the device is recognized again within the first setting time. After the first setting time is reached, if the signal recognition result is still not received or no target action is recognized within the first setting time, the device turns off the first timer and updates the device processing result, increasing the historical failure result represented by the device processing result. For example, increment the historical failure count (Fail) in the historical failure result by one.
[0102] For example, after the device sends the first voice signal to the gateway, it can start a timer and set a timing duration of 3s. Within these 3s, if the device receives feedback from the gateway that the first voice signal is a valid voice command, then increment the historical success count Successful by 1, initialize the historical failure count Fail to 0, and continue to wait for whether a new action is triggered or wait for a timeout interruption after 3s of timing.
[0103] In this embodiment, after the first voice signal is sent, the first timer is set; then the device continuously waits for the target device to feedback the signal recognition result of the first voice signal within the first setting time. If the signal recognition result obtained by the device indicates that the first voice signal is valid, the device processing result is initialized and the historical success count is increased. By setting the first timer, the waiting time of the device is controlled, reducing the energy consumption of the device during the waiting process.
[0104] In one embodiment, the first timer is further configured to indicate that within a first set time, action recognition processing is performed on the acquired current action acquisition information to obtain a current action recognition result. After setting the first timer, the following steps are further included: within the first set time of the first timer, if it is detected that the current action represented by the current action recognition result matches the target action, then a second voice signal is collected; the second voice signal is sent to the target device; and a signal recognition result for the second voice signal returned by the target device is obtained.
[0105] Wherein, the second voice signal refers to the voice signal newly collected by the target device within the first set time after successful networking.
[0106] Specifically, after the device finishes sending the first voice signal, within the first time of the first timer, in addition to waiting for the target device to feedback the signal recognition result of the previous voice signal (for example), the device can also turn on the power of the inertial sensor and the interrupt of the acceleration sensor in the inertial sensor to continuously detect the action acquisition information through the acceleration sensor, that is, obtain the current action acquisition information; the device performs action recognition processing on the current action acquisition information to identify the action currently triggered by the device, and thus obtains the current action recognition result. The current action recognition result is matched with the target action; if it is detected that the current action represented by the current action recognition result matches the target action, indicating that the user wants to provide a new voice command, then the device collects a voice signal again to obtain a second voice signal. The device sends the second voice signal to the target device; then the device can reset the timing of the first set time corresponding to the first timer, and the device continues to wait for the target device to feedback the signal recognition result for the second voice signal within the reset first set time.
[0107] Furthermore, the user can also trigger the device again through the target action within the first set time after sending the second voice signal, so that the device collects a third voice signal. Then the device will reset the timing of the first set time corresponding to the first timer again, and the device will execute the steps of transmitting the third voice signal and waiting for the target device to feedback the signal recognition result for the newly sent voice signal again. Similarly, the user can also transmit a fourth voice signal, a fifth voice signal, a sixth voice signal, and so on by triggering the target action within the first set time. It should be noted that the newly sent voice signal can be a brand-new command, and this command can overwrite the old command or be other action commands, depending on the specific needs of the user.
[0108] For example, the device sends the voice signal (1) to the gateway; after the sending is completed, it turns on the interrupt of the 6 IMU power supply and the ACC, starts the timer and sets it to 3 seconds; within these 3 seconds, the device continuously reads the ACC value and identifies whether there is a gesture of "picking up" or "shaking" based on the ACC value, while waiting for the gateway to feedback whether the previous voice signal (1) is recognized as a valid instruction. If the device recognizes a gesture of "picking up" or "shaking" within these 3 seconds, the device stops reading the ACC value, turns off the ACC interrupt and the power supply of the 6 IMU, the device collects the voice signal (2) again and returns to the step of sending the voice signal (2) to the gateway; if the device receives the feedback from the gateway that the previous voice signal is a valid instruction, then Successful is incremented by 1 and Fail is set to 0, and it continues to wait for whether a gesture of "picking up" or "shaking" is recognized or for the 3 - second period to end. After 3 seconds, if neither a gesture of "picking up" or "shaking" is recognized nor the gateway feedback on whether the previous voice signal is recognized as a valid instruction is received, the device turns off the timer, turns off the Bluetooth, increments Fail by 1 and then returns to the sleep state.
[0109] In this embodiment, if within the first set time of the first timer, it is detected that the current action represented by the current action recognition result matches the target action, then a new voice signal is re - collected and sent to the target device; then the first timer is reset, and within a new round of the first set time, it continues to wait for the signal recognition result of the newly sent voice signal returned by the target device, achieving the collection and rapid response of the user's new instruction. While saving the device power consumption, it also improves the response speed of the device.
[0110] In one embodiment, after setting the first timer, it further includes: when the first set time of the first timer is reached, if the current action represented by the current action recognition result does not match the target action and no signal recognition result indicating the validity of the first voice signal is obtained, then the device processing result is updated and the network connection processing is ended.
[0111] Specifically, when the first set time of the first timer is reached, if neither an action matching the target action is recognized nor a signal recognition result indicating the validity of the first voice signal is obtained, then the device can turn off the first timer, turn off the Bluetooth, end the network connection processing, and update the device processing result, increasing the historical failure result represented by the device processing result, for example, incrementing the historical failure count (Fail) in the historical failure result by 1. Finally, the device enters the sleep state.
[0112] In this embodiment, when the first set time of the first timer is reached, if the current action matching the target action has not been recognized yet and no signal recognition result indicating that the feedback voice signal is valid has been received, the device processing result is updated and the network connection process is ended to reduce energy consumption and effectively reduce the power consumption of the device.
[0113] In one embodiment, the above device control method further includes: setting a second timer; when the second set time corresponding to the second timer is reached, obtaining an interruption generated by the second timer; the interruption is used to indicate initializing the device processing result.
[0114] Wherein, the second timer is a timer used to indicate that the device enters the sleep state. For example, the second timer can be a sleep timer.
[0115] Specifically, the device can set the second timer and the second set time corresponding to the second timer. If the second timer reaches the second set time, the second timer generates an interruption to wake up the device. The device turns off the second timer and initializes both the historical failure times and the historical success times in the historical failure result represented by the device processing result, and then the device enters the sleep state. For example, if the network connection fails, the device can set a 5-minute sleep timer. After 5 minutes, the device initializes both Successful and Fail to 0, and then the device enters the sleep state.
[0116] It should be noted that since the historical failure times are initialized, the next time the device is triggered, if the device processing result does not exceed the second preset threshold, the network connection process can be directly executed to obtain the network connection result without performing action recognition.
[0117] In this embodiment, by setting the second timer, the device processing result of the device can be initialized periodically, so that the device can quickly enter the network connection process without performing action recognition when it is triggered next time after initialization, greatly improving the network connection speed of the device and the speed of responding to user instructions.
[0118] In one embodiment, as Figure 4 shown, another device control method is provided. Taking the method applied to the device in Figure 1 as an example, it includes the following steps:
[0119] Step S401, if the action acquisition information of the device exceeds the first preset threshold, obtain the device processing result.
[0120] According to whether the device processing result exceeds the second preset threshold, select to execute step S402 or step S406.
[0121] Step S402, if the device processing result exceeds the second preset threshold, perform action recognition processing based on the action acquisition information to obtain an action recognition result; the action recognition result is used to indicate whether to perform network connection processing.
[0122] According to whether the action recognition result matches the target action, select to execute Step S403 or Step S404.
[0123] Step S403, if the recognized action represented by the action recognition result matches the target action, collect the first voice signal and perform network connection processing with the target device to obtain a network connection result with the target device.
[0124] Step S404, if the recognized action represented by the action recognition result does not match the target action, update the device processing result.
[0125] Step S405, set a second timer; when the second set time corresponding to the second timer is reached, obtain the interrupt generated by the second timer; the interrupt is used to indicate initializing the device processing result.
[0126] Step S406, if the device processing result does not exceed the second preset threshold, collect the first voice signal and perform network connection processing to obtain a network connection result.
[0127] According to whether the network connection is successful, select to execute Step S407 or Step S408.
[0128] Step S407, if the network connection result indicates successful network connection, send the first voice signal to the target device so that the target device performs the action corresponding to the first voice signal.
[0129] Step S408, if the network connection result indicates a failed network connection, update the device processing result. Jump to Step S405.
[0130] The above device control method can achieve the following beneficial effects: By combining the action acquisition information of the device and the device processing result for judgment and decision-making, if the device processing result exceeds the threshold, it indicates that there is a high failure rate when the device performs the target processing. At this time, action recognition processing is preferentially performed based on the action acquisition information to reduce unnecessary network connection processing and energy consumption. If the device processing result does not exceed the threshold, network connection processing is directly performed to provide fast response. Through this dynamic control mechanism, both the response timeliness of the device is improved and the power consumption of the device is reduced, thus effectively balancing the power consumption of the device and the response speed of network connection processing.
[0131] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0132] Based on the same inventive concept, an embodiment of the present application further provides a device control device for implementing the above-mentioned device control method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device control device provided below can refer to the limitations on the device control method in the above text, and will not be repeated here.
[0133] In one embodiment, as Figure 5 shown, a device control device 500 is provided, including: an action measurement module 501, an action recognition module 502, and a networking processing module 503, where:
[0134] The action measurement module 501 is configured to obtain the current device processing result if the action acquisition information of the device exceeds a first preset threshold; the action acquisition information is used to represent the sensor data measured when the device is triggered; the device processing result is used to represent the historical failure result of the device performing the target processing; the target processing includes at least one of action recognition processing and voice processing.
[0135] The action recognition module 502 is configured to perform action recognition processing based on the action acquisition information to obtain an action recognition result if the device processing result exceeds a second preset threshold; the action recognition result is used to indicate whether to perform networking processing.
[0136] The networking processing module 503 is configured to perform networking processing to obtain a networking result if the device processing result does not exceed the second preset threshold.
[0137] In one embodiment, the device control device 500 further includes a result judgment module, configured to perform networking processing to obtain a networking result if the recognition action represented by the action recognition result matches the target action; and update the device processing result if the recognition action represented by the action recognition result does not match the target action.
[0138] In one embodiment, the networking processing module 503 is further configured to collect a first voice signal and perform networking processing with a target device to obtain a networking result with the target device.
[0139] In one embodiment, the device control device 500 further includes a voice transmission module, configured to: if the networking result indicates successful networking, send the first voice signal to the target device so that the target device can recognize the first voice signal; if the networking result indicates failed networking, update the device processing result.
[0140] In one embodiment, the device control device 500 further includes a first timing module, configured to set a first timer; within a first set time corresponding to the first timer, if a signal recognition result for the first voice signal returned by the target device is obtained and the signal recognition result indicates that the first voice signal is valid, initialize the device processing result.
[0141] In one embodiment, the first timer is further configured to indicate that within the first set time, perform action recognition processing on the acquired current action acquisition information to obtain a current action recognition result. The device control device 500 further includes an action detection module, configured to: within the first set time of the first timer, if it is detected that the current action indicated by the current action recognition result matches the target action, collect a second voice signal; send the second voice signal to the target device so that the target device can recognize the second voice signal.
[0142] In one embodiment, the device control device 500 further includes a networking end module, configured to: when the first set time of the first timer is reached, if the current action indicated by the current action recognition result does not match the target action and a signal recognition result indicating that the first voice signal is valid is not obtained, update the device processing result and end the networking processing.
[0143] In one embodiment, the device control device 500 further includes a second timing module, configured to set a second timer; when the second set time corresponding to the second timer is reached, obtain an interruption generated by the second timer; the interruption is used to indicate initializing the device processing result.
[0144] Each module in the above device control device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the controller in the voice processing device in hardware form or independent of it, or stored in the memory in the voice processing device in software form, so that the controller can call and execute the operations corresponding to the above modules.
[0145] In one embodiment, a voice processing device is provided. The voice processing device can be a voice companion, and its internal structure diagram can be as Figure 6As shown. The voice processing device includes a controller, a sensor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the controller, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the controller of the voice processing device is used to provide computing and control capabilities. The memory of the voice processing device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the voice processing device is used to store data such as action acquisition information, action recognition results, and networking results. The input / output interface of the voice processing device is used to exchange information between the controller and external devices. The communication interface of the voice processing device is used to communicate with an external terminal through a network connection. When the computer program is executed by the controller, it implements a device control method.
[0146] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the voice processing device to which the solution of this application is applied. The specific voice processing device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0147] In one embodiment, a voice processing device is further provided, including a memory, a sensor, and a controller. A computer program is stored in the memory. When the controller executes the computer program, the steps in the above method embodiments are implemented.
[0148] The sensor is used to obtain action measurement value acquisition information of the voice processing device. If the action acquisition information exceeds the first preset threshold of the action threshold, it instructs the controller to obtain the current device processing result; sends the action measurement value to the controller; the action measurement value acquisition information is used to characterize the sensor data measured when the voice processing device is triggered; the device processing result is used to characterize the historical failure result of the voice processing device performing the target processing; the target processing includes at least one of action recognition processing and voice processing.
[0149] The controller is used to perform action recognition processing on the received action acquisition information to obtain an action recognition result if it detects that the device processing result exceeds the second preset threshold; the action recognition result is used to indicate whether to perform networking processing; if the device processing result does not exceed the second preset threshold, perform networking processing to obtain a networking result.
[0150] Among them, the voice processing device can be a voice remote control. The voice remote control can also be connected to a target device (such as a gateway) through Bluetooth.
[0151] Figure 7 It is a schematic diagram of the hardware system of a voice processing device. The voice processing device can be powered by a detachable button battery CR2450 for the entire voice processing device. The main control MG24 controls the computing, control, communication, etc. of the entire voice processing device. A six-axis inertial sensor (6 IMU) transmits measurement data to the main control MG24 through the IIC interface. The main control MG24 performs motion recognition on the signals of the 6 IMU. Generally, the 6 IMU works in the low-power mode. Two digital microphones (MIC) are used as voice collection units to collect voice signals and transmit the collected voice signals to the main control MG24 through the IIS interface. They are connected to the same group of IIS and distinguish who is transmitting signals on the bus through the signals of the WS pins. There is also a buzzer, three RGB indicator lights, a Nor FLASH memory, and a reset button in the voice processing device.
[0152] Specifically, the sensor in the voice processing device is responsible for obtaining the action measurement value acquisition information of the voice processing device; the controller in the voice processing device is responsible for device control based on the action acquisition information. For the specific processing process, refer to the above device control method and will not be elaborated here.
[0153] In this embodiment, the action acquisition information of the voice processing device is collected through the sensor. If the action acquisition information exceeds the first preset threshold of the action threshold, the controller is instructed to obtain the current device processing result. Then, the controller detects whether the action acquisition information sent by the sensor exceeds the second preset threshold, and performs action recognition processing on the action acquisition information to obtain an action recognition result when the device processing result exceeds the second preset threshold. If the device processing result does not exceed the second preset threshold, network connection processing is performed to obtain a network connection result. By combining the action acquisition information of the sensor and the device processing result of the controller for judgment and decision-making, if the device processing result exceeds the threshold, it indicates that there is a high failure rate when the device performs the target processing. At this time, action recognition processing is preferentially performed based on the action acquisition information to reduce unnecessary network connection processing and energy consumption. If the device processing result does not exceed the threshold, network connection processing is directly performed to provide a quick response. Through this dynamic control mechanism, both the response timeliness of the device is improved and the power consumption of the voice processing device is reduced, thus effectively balancing the power consumption of the voice processing device and the response speed of the network connection processing.
[0154] In one embodiment, the controller is further configured to, after the voice processing device is started, control the sensor to enter a first working mode and cause the controller to enter a second working mode; the power consumption of the second working mode is lower than that of the first working mode. The sensor is further configured to obtain motion acquisition information of the voice processing device in the first working mode, and when the motion acquisition information exceeds a first preset threshold, wake up the controller from the second working mode and instruct the controller to obtain a device processing result.
[0155] Wherein, the first working mode may be a low-power mode, and the second working mode may be a sleep state.
[0156] Specifically, the voice processing device performs initialization and parameter setting: when the voice processing device is started, the historical success times Successful and the historical failure times Fail are initialized to 0, and the historical success times and the historical failure times are used to track the success and failure times of voice transmission and motion recognition. The voice processing device enters the low-power mode: the main control MG24 configures the interruption threshold of 6 IMUs, enables the 6 IMUs to work in the low-power mode, and causes the controller to enter the sleep state to save energy. When the 6 IMUs detect motion acquisition information exceeding the interruption threshold, the main control MG24 is woken up. MG24 decides whether to directly perform motion recognition according to the historical failure times Fail.
[0157] In this embodiment, waking up the controller for device control only when the motion acquisition information exceeds the interruption threshold can reduce the power consumption loss caused by accidental touch, thereby increasing the service life of the voice processing device, and at the same time, the voice processing device can be quickly woken up when voice control is required, improving the usage fluency of the voice processing device.
[0158] In one embodiment, the voice processing device further includes a wireless communication unit. The controller is further configured to, if the recognized motion represented by the motion recognition result matches the target motion, control the wireless communication unit to perform networking processing to obtain a networking result; if the recognized motion represented by the motion recognition result does not match the target motion, increment the device processing result by one and enter the second working mode again according to the set timing information.
[0159] Specifically, if the historical failure times Fail are greater than 5, the main control MG24 reads the recognized motion represented by the motion acquisition information of the 6 IMUs and determines whether the recognized motion matches the target motion ("pick up" or "shake"). If the two match, the networking process is entered; if the two do not match, a 5-minute sleep timer is set, the historical failure times Fail are incremented by 1, the historical success times Successful are set to 0, and then the voice processing device returns to the sleep state.
[0160] In this embodiment, while maintaining low power consumption, the voice processing device can also determine whether to immediately attempt to connect to the network or perform action recognition based on the historical failure count, achieving the dual goals of fast response and extended device lifespan.
[0161] To more clearly illustrate the device control method provided by the embodiments of the present disclosure, the above device control method will be specifically described below with a specific embodiment. As Figure 8 shown, another device control method is provided, which can be applied to the Figure 1 device, and specifically includes the following content:
[0162] First step: After the device enters normal operation, it will set two parameters, Successful and Fail, to mark the number of successful and failed executions of the language signal transmission and recognition actions; the initial values of Successful and Fail are both set to 0; the main control MG24 turns on the power of the 6 IMU, configures the interruption threshold of its internal ACC, and enables the ACC interruption (when the value measured by the ACC in the low power consumption mode exceeds the interruption threshold, the 6 IMU outputs an interruption signal to wake up MG24).
[0163] Second step: The device enters the sleep state: the main control MG24 and Nor FLASH are in the sleep state, the ACC in the 6 IMU is in the low power consumption operation mode, and the power supplies of other peripherals are all turned off. The device is in the sleep state without being moved at all.
[0164] Third step: When the device is moved and the measured value of the ACC exceeds the set interruption threshold, the 6 IMU outputs an interruption signal to wake up MG24, and the device is woken up (the main control MG24 is woken up) and MG24 clears and turns off the sleep timer; then MG24 determines whether Fail is greater than 5: If Fail > 5, then the main control MG24 reads the value of the ACC and identifies whether the device is "picked up" or "shaken" at this time; if it is not "picked up" or "shaken", then MG24 turns on the internal sleep timer and sets it to 5 minutes, while Fail is incremented by 1 and Successful is set to 0; then the device enters the second step; if it is "picked up" or "shaken", then the device enters the fourth step; if Fail ≤ 5, then the device enters the fourth step.
[0165] Fourth step: MG24 turns off the ACC interruption and the power of the 6 IMU, starts to initiate the BLE link to the gateway (connect to the network), and synchronously collects the language signal; if the network connection fails, then MG24 turns on the internal sleep timer and sets it to 5 minutes, while Fail is incremented by 1 and Successful is set to 0; then the device enters the second step; if the network connection is successful, then the device enters the fifth step.
[0166] Step 5: Start sending the voice signal to the gateway; after the sending is completed (MG24 sets the maximum voice sending duration to 10 s, and if it exceeds 10 s, it will automatically stop sending and consider the sending completed), turn on the power of the 6 IMUs and the interrupt of the ACC, and start the timer and set it to 3 seconds.
[0167] Step 6: Within these 3 s, MG24 continuously reads the value of the ACC and calculates and identifies whether there are gestures of "pick up" or "shake", and waits for the gateway to feedback whether the previous voice signal is recognized as a valid instruction; if the gestures of "pick up" or "shake" are recognized, MG24 stops reading the data of the ACC, turns off the ACC interrupt and the power of the 6 IMUs, MG24 collects the voice signal again and returns to Step 5; if the feedback received is "valid instruction", Successful is incremented by 1 and Fail is set to 0 and continues to wait.
[0168] Step 7: If neither the gesture of "pick up" or "shake" is recognized nor the gateway feedbacks whether the previous voice signal is recognized as a valid instruction, MG24 turns off the timer, turns off the BLE, and increments Fail by 1 and then returns to Step 2.
[0169] Step 8: If the set sleep timer reaches the set time, the sleep timer generates an interrupt to wake up MG24. MG24 turns off the sleep timer and sets Successful and Fail to 0, and then enters Step 2.
[0170] In this embodiment, by combining the action collection information of the sensor and the device processing result of the controller for judgment and decision-making, the power consumption loss caused by accidental touch can be reduced, thereby increasing the service life of the voice processing device. If the device processing result exceeds the threshold, it indicates that there is a high failure rate when the device performs the target processing. At this time, the action recognition processing is preferentially based on the action collection information to reduce unnecessary networking processing and energy consumption. If the device processing result does not exceed the threshold, the networking processing is directly executed to provide a quick response. Through this dynamic control mechanism, both the response timeliness of the device is improved and the power consumption of the voice processing device is reduced, thus effectively balancing the power consumption of the voice processing device and the response speed of the networking processing.
[0171] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the controller, the steps in the above method embodiments are implemented.
[0172] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the controller, the steps in the above method embodiments are implemented.
[0173] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The controllers involved in the embodiments provided in this application can be general controllers, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0174] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0175] The above-described embodiments only represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A device control method, characterized in that: The method comprises: If the acquired motion collection information of the device exceeds a first preset threshold, the device processing result is acquired; the motion collection information is used to characterize the sensor data measured when the device is triggered; the device processing result is used to characterize the historical failure result of the device in executing the target processing; the target processing includes at least one of motion recognition processing and voice processing; If the device processing result exceeds the second preset threshold, the action recognition processing is performed based on the action collection information to obtain the action recognition result; the action recognition result is used to indicate whether to perform network processing; If the device processing result does not exceed the second preset threshold, the networking processing is performed to obtain a networking result.
2. The method according to claim 1, characterized in that After performing the action recognition process based on the action collection information to obtain the action recognition result, the method further includes: If the recognized action represented by the action recognition result matches the target action, the networking process is performed to obtain the networking result; If the recognized action represented by the action recognition result does not match the target action, the device processing result is updated.
3. The method according to claim 1, characterized in that: The performing of the networking process to obtain a networking result includes: A first voice signal is collected, and networking processing is performed with a target device to obtain the networking result between the target device and the target device.
4. The method according to claim 3, characterized in that After executing the networking process and obtaining the networking result, the method further includes: If the networking result indicates that the networking is successful, sending the first voice signal to the target device so that the target device recognizes the first voice signal; If the networking result indicates that the networking fails, the device processing result is updated.
5. The method according to claim 4, characterized in that After sending the first voice signal to the target device, the method further includes: Setting a first timer; Within a first set time corresponding to the first timer, if a signal recognition result for the first voice signal returned by the target device is obtained, and the signal recognition result indicates that the first voice signal is valid, the device processing result is initialized.
6. The method according to claim 5, characterized in that The first timer is also used to instruct to perform action recognition processing on the acquired current action collection information within the first set time to obtain a current action recognition result; After setting the first timer, the method further includes: Within the first setting time of the first timer, if it is detected that the current action represented by the current action recognition result matches the target action, collecting a second voice signal; sending the second voice signal to the target device so that the target device recognizes the second voice signal; Acquire a signal recognition result for the second voice signal returned by the target device.
7. The method according to claim 6, characterized in that After setting the first timer, the method further includes: When the first set time of the first timer is reached, if the current action represented by the current action recognition result does not match the target action, and the signal recognition result representing that the first voice signal is valid is not obtained, the device processing result is updated and the networking processing is terminated.
8. The method according to any one of claims 2 or 4, characterized in that: The method further comprises: Setting a second timer; When a second set time corresponding to the second timer is reached, an interrupt generated by the second timer is obtained; the interrupt is used to instruct to initialize the device processing result.
9. A device control device, characterized in that: The device comprises: An action measurement module, configured to obtain a current device processing result if the acquired device action collection information exceeds a first preset threshold; the action collection information is used to characterize the sensor data measured when the device is triggered; the device processing result is used to characterize the historical failure result of the device in executing the target processing; the target processing includes at least one of action recognition processing and voice processing; an action recognition module, configured to perform the action recognition processing based on the action collection information to obtain an action recognition result if the device processing result exceeds a second preset threshold; the action recognition result is used to indicate whether to perform network processing; The networking processing module is used to execute the networking processing to obtain the networking result if the device processing result does not exceed the second preset threshold.
10. A speech processing device, characterized in that: The speech processing device includes a sensor and a controller; The sensor is used to obtain the action collection information of the voice processing device, and if the action collection information exceeds a first preset threshold, instruct the controller to obtain the current device processing result; the action collection information is used to represent the sensor data measured when the voice processing device is triggered; The device processing result is used to characterize the historical failure result of the speech processing device in executing the target processing; the target processing includes at least one of action recognition processing and speech processing; The controller is used to perform the motion recognition processing on the received motion collection information to obtain the motion recognition result if it is detected that the device processing result exceeds the second preset threshold; the motion recognition result is used to indicate whether to perform networking processing; If the device processing result does not exceed the second preset threshold, the networking processing is performed to obtain a networking result.
11. The speech processing device according to claim 10, characterized in that: The controller is further used to control the sensor to enter a first working mode and the controller to enter a second working mode after the voice processing device is started; the power consumption of the second working mode is lower than that of the first working mode; The sensor is also used to obtain the motion collection information of the voice processing device in the first working mode, and when the motion collection information exceeds the first preset threshold, wake up the controller from the second working mode and instruct the controller to obtain the device processing result.
12. The speech processing device according to claim 11, characterized in that: The speech processing device also includes a wireless communication unit; The controller is also used to control the wireless communication unit to perform the networking processing to obtain the networking result if the recognition action represented by the action recognition result matches the target action; if the recognition action represented by the action recognition result does not match the target action, add one to the device processing result and enter the second working mode again according to the set timing information.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a controller, the steps of the method according to any one of claims 1 to 8 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a controller, the steps of the method according to any one of claims 1 to 8 are implemented.