Equipment control method and device, smart home equipment and storage medium

By acquiring environmental information and historical voice recognition information of smart home devices, a target score is calculated to control their continuous speaking mode. This solves the problem of inaccurate control caused by single environmental noise, improves the accuracy and reliability of the continuous speaking mode, and optimizes the user experience.

CN121963723APending Publication Date: 2026-05-01GREE ELECTRIC APPLIANCE INC OF ZHUHAI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREE ELECTRIC APPLIANCE INC OF ZHUHAI
Filing Date
2025-12-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the continuous mode control of smart home devices relies on a single ambient noise, which leads to inaccurate control and affects the user experience.

Method used

By acquiring environmental information and historical speech recognition information of the target device, the correlation between the target speech command and the historical speech command set is determined. The target score is calculated by combining the environmental information, historical speech recognition information and correlation, and is used to control the continuous speaking mode of the target device.

Benefits of technology

It achieves multi-dimensional comprehensive control over continuous speaking mode, adapts to environmental changes and voice interaction situations, improves the accuracy and reliability of control, and ensures user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963723A_ABST
    Figure CN121963723A_ABST
Patent Text Reader

Abstract

The invention relates to an equipment control method and device, smart home equipment and a storage medium, and the method comprises the steps: obtaining the environment information of the environment where the current target equipment is located and the historical voice recognition information corresponding to the target equipment when the target equipment receives a target voice instruction; based on the target voice instruction, determining a target association degree between the target voice instruction and a historical voice instruction set received by the target device; based on the environment information, the historical voice recognition information and the target association degree, determining a target score corresponding to the target device, the target score being used for indicating and determining a control strategy for controlling a continuous speaking mode of the target device; and controlling the target equipment to execute a control strategy corresponding to the target score. According to the invention, the accuracy and reliability of continuous speaking mode control are improved, and the user experience is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

A method for controlling a device, an apparatus, a smart home device, and a storage medium. Technical Field

[0001] This application relates to the field of voice interaction technology, and in particular to a control method, apparatus, smart home device, and storage medium for a device. Background Technology

[0002] With the widespread adoption of voice interaction technology in various smart home devices, users often need to issue multiple related voice commands consecutively. If each voice command requires waking up the smart home device, it significantly increases the complexity of operation. Therefore, a continuous voice command mode for smart home devices has emerged, allowing the device to receive multiple consecutive voice commands without repeated wake-up. Currently, controlling the continuous voice command mode of smart home devices typically relies on a single ambient noise source. However, using only ambient noise to control the continuous voice command mode does not consider the influence of other factors, leading to inaccuracies and impacting the user experience. Summary of the Invention

[0003] This application provides a device control method, apparatus, smart home device, and storage medium to solve the problem that the control of continuous mode in smart home devices in the prior art is inaccurate, which affects the user experience.

[0004] In a first aspect, this application provides a device control method, comprising: when a target device receives a target voice command, acquiring environmental information of the current environment in which the target device is located and historical speech recognition information corresponding to the target device; determining a target correlation degree between the target voice command and a set of historical voice commands received by the target device based on the target voice command; determining a target score corresponding to the target device based on the environmental information, the historical speech recognition information and the target correlation degree, wherein the target score is used to indicate a control strategy for controlling the continuous speaking mode of the target device; and controlling the target device to execute the control strategy corresponding to the target score.

[0005] In an optional implementation, determining the target score corresponding to the target device based on the environmental information, the historical speech recognition information, and the target relevance includes: determining a first weight corresponding to the environmental information, a second weight corresponding to the historical speech recognition information, and a third weight corresponding to the target relevance; and using the first weight, the second weight, and the third weight, performing a weighted summation on the environmental information, the historical speech recognition information, and the target relevance to obtain the target score corresponding to the target device.

[0006] In an optional implementation, the environmental information includes environmental noise level and number of environmental objects, and the first weight includes a first sub-weight corresponding to the environmental noise level and a second sub-weight corresponding to the number of environmental objects; determining the first weight corresponding to the environmental information, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target correlation includes: obtaining the initial weights corresponding to the environmental noise level, the number of environmental objects, the historical speech recognition information, and the target correlation; determining the first number of times the target device has erroneously triggered a voice event within a first preset time period and the second number of times the target device has successfully recognized a voice command consecutively before the current moment; and correcting each of the initial weights according to the first number and the second number to obtain the first sub-weight corresponding to the environmental noise level, the second sub-weight corresponding to the number of environmental objects, the second sub-weight corresponding to the historical speech recognition information, and the third sub-weight corresponding to the target correlation.

[0007] In an optional implementation, all historical voice commands in the historical voice command set are received by the target device within a second preset time period and are all successfully recognized by the target device; determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the target voice command includes: determining a target historical voice command from the historical voice command set, wherein the target historical voice command is the historical voice command preceding the target voice command; determining a first semantic correlation degree between the target voice command and the target historical voice command; determining a second logical correlation degree between the target voice command and the historical voice command set; and determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the first correlation degree and the second correlation degree.

[0008] In an optional implementation, determining the second logical correlation between the target voice command and the historical voice command set includes: extracting features from the target voice command and each historical voice command in the historical voice command set based on a preset feature dimension set to obtain target features of the target voice command under each preset feature dimension in the preset feature dimension set and historical features of each historical voice command in the historical voice command set under each preset feature dimension in the preset feature dimension set; determining the logical matching degree between the target voice command and each historical voice command in the historical voice command set based on all target features corresponding to the target voice command and all historical features corresponding to each historical voice command in the historical voice command set; determining the target weight of the logical matching degree corresponding to each historical voice command based on the historical reception time corresponding to each historical voice command in the historical voice command set, wherein the target weight decreases as the time distance between the historical reception time and the current time increases; and weighted summing all the obtained logical matching degrees using all the obtained target weights to obtain the second logical correlation between the target voice command and the historical voice command set.

[0009] In an optional implementation, determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the first correlation degree and the second correlation degree includes: determining a first comparison result between the first correlation degree and a first correlation degree threshold, and a second comparison result between the second correlation degree and a second correlation degree threshold; determining a fourth weight corresponding to the first correlation degree and a fifth weight corresponding to the second correlation degree based on the first comparison result and the second comparison result; and performing a weighted summation of the first correlation degree and the second correlation degree using the fourth weight and the fifth weight to obtain the target correlation degree between the target voice command and the historical voice command set received by the target device.

[0010] In an optional implementation, controlling the target device to execute the control strategy corresponding to the target score includes: when the target score is greater than or equal to a first score threshold, controlling the target device to operate in the continuous speaking mode for a first preset duration; when the target score is less than a second score threshold, controlling the continuous speaking mode of the target device to be turned off, wherein the second score threshold is less than the first score threshold; and when the target score is greater than or equal to the second score threshold and less than the first score threshold, controlling the target device to operate in the continuous speaking mode for a second preset duration, wherein the second preset duration is less than the first preset duration.

[0011] Secondly, this application provides a control device for a device, comprising: an acquisition module, configured to acquire environmental information of the current environment of the target device and historical speech recognition information corresponding to the target device when the target device receives a target speech command; a determination module, configured to determine a target correlation degree between the target speech command and a set of historical speech commands received by the target device based on the target speech command; the determination module is further configured to determine a target score corresponding to the target device based on the environmental information, the historical speech recognition information and the target correlation degree, the target score being used to indicate a control strategy for controlling the continuous speaking mode of the target device; and a control module, configured to control the target device to execute the control strategy corresponding to the target score.

[0012] Thirdly, this application provides a smart home device, including: a processor and a memory, wherein the processor is used to execute a device control program stored in the memory to implement the device control method described above.

[0013] Fourthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the device control method described above.

[0014] Compared with the prior art, the technical solutions provided in this application have the following advantages. The device control method provided in this application includes: when the target device receives a target voice command, acquiring environmental information of the current environment of the target device and historical speech recognition information corresponding to the target device; determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the target voice command; determining the target score corresponding to the target device based on the environmental information, historical speech recognition information and target correlation degree, wherein the target score is used to indicate the control strategy for controlling the continuous speaking mode of the target device; and controlling the target device to execute the control strategy corresponding to the target score. Through the above methods, this embodiment simultaneously acquires the current environmental information of the target device and the corresponding historical speech recognition information when the target device receives the target voice command. Simultaneously, it determines the target correlation degree between the target voice command and the historical voice command set received by the target device. By fusing environmental information, historical speech recognition information, and target correlation degree, it obtains the target score corresponding to the target device. Based on the control strategy determined by this target score, it controls the continuous speaking mode of the target device, achieving multi-dimensional comprehensive management and control of the continuous speaking mode. This allows for simultaneous adaptation to environmental changes and the voice interaction of the target device, avoiding the inaccurate control of the continuous speaking mode caused by relying solely on single environmental noise. This improves the accuracy and reliability of the continuous speaking mode control, ensuring a superior user experience. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0018] Figure 1 is a flowchart illustrating a device control method according to an embodiment of this application; Figure 2 is a flowchart illustrating a device control method according to an embodiment of this application; Figure 3 is a structural diagram illustrating a device control apparatus according to an embodiment of this application; Figure 4 is a structural diagram illustrating a smart home device according to an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0021] Referring to Figure 1, Figure 1 is a flowchart illustrating a device control method provided in an embodiment of this application. The device control method provided in this embodiment includes the following steps: S101: When the target device receives a target voice command, it acquires environmental information of the current environment of the target device and historical voice recognition information corresponding to the target device.

[0022] In this embodiment, the method is applied to a target device, which can be a smart home device, including smart air conditioners, smart washing machines, and smart speakers. The target voice command is the voice command received by the target device after the user wakes it up. Environmental information includes the environmental noise level and the number of people in the environment. The environmental noise level can be divided into low-level noise, medium-level noise, and high-level noise, and the number of people in the environment can be divided into single and multiple people. The environmental noise level can be obtained by collecting the ambient sound signal of the environment in which the target device is located through a microphone array in the target device, and then analyzing the sound signal to obtain the environmental noise level. The number of people in the environment can be obtained by detecting the sound signal through millimeter-wave radar in the target device.

[0023] Historical speech recognition information can be understood as records of the target device's recognition of received voice commands within a third preset time period. This third preset time period is prior to the current moment and is temporally continuous with the current moment. The historical speech recognition information can be the success rate of the target device recognizing received voice commands within the third preset time period. This success rate can be determined by statistically analyzing the total number of first-level voice commands successfully recognized by the target device within the third time period and the total number of second-level voice commands received by the target device within the third time period. The ratio between the total number of first-level and second-level voice commands is then used to determine the success rate of the target device recognizing received voice commands within the third preset time period. The third preset time period in the above formula can be set according to actual needs; this embodiment does not limit it.

[0024] Historical speech recognition information can also be used to determine the failure rate of the target device in recognizing the received voice commands within the third preset time period. This can be achieved by statistically analyzing the number of interruptions by the user within the third preset time period and the total number of third commands received by the target device within the third time period. The ratio between the number of interruptions and the total number of third commands can be used to determine the failure rate of the target device in recognizing the received voice commands within the third preset time period.

[0025] S102: Based on the target voice command, determine the target correlation degree between the target voice command and the historical voice command set received by the target device.

[0026] In this embodiment, the historical voice command set includes multiple historical voice commands, which are voice commands received by the target device before the current moment. To ensure the accuracy of subsequent control of the continuous speaking mode, the historical voice commands are actually voice commands successfully recognized by the target device. The historical voice command set can consist of all historical voice commands received by the target device within a second preset time period. The second preset time period is located before the current moment and is temporally continuous with the current moment. To ensure data consistency, the second preset time period can be consistent with a third preset time period. Target correlation is used to quantify the coherence between the target voice command and the historical voice command set, and its function is to determine whether the current target voice command is a continuation of the original dialogue.

[0027] After acquiring the historical voice command set within the second preset time period, the target voice command and the historical voice command set are analyzed for coherence to obtain the target correlation degree between the target voice command and the historical voice command set detected by the target device. Based on the target correlation degree, combined with the current environmental information and historical voice recognition information, the continuous speaking mode control of the target device is realized.

[0028] S103: Determine the target score corresponding to the target device based on environmental information, historical speech recognition information, and target correlation.

[0029] In this embodiment, the target score is used to indicate the control strategy for controlling the continuous speaking mode of the target device. The target score is calculated by fusing environmental information, historical speech recognition information, and target correlation using a built-in preset algorithm. After obtaining the target score, a control strategy corresponding to the target score can be determined. This control strategy characterizes the control method for controlling the continuous speaking mode of the target device. For example, if the control strategy is to control the target device to operate in continuous speaking mode, then the control operation corresponding to the above control strategy is executed so that the target device operates in continuous speaking mode.

[0030] The preset algorithm can be obtained by weighted summation of environmental information, historical speech recognition information, and target relevance, or it can be obtained by decision-making using a pre-trained decision tree model. The weighted summation algorithm will be described below and will not be repeated here. Decision-making using a decision tree model can be achieved as follows: The decision tree model includes multiple leaf nodes, each corresponding to node features and a node score. The node features include environmental information, historical speech recognition information, and target relevance. After obtaining the environmental information, historical speech recognition information, and target relevance, the target node to which the environmental information, historical speech recognition information, and target relevance belong is determined. The node score corresponding to that target node is determined as the target score for the target device, thus achieving fast and accurate determination of the target score through the decision tree model.

[0031] S104: Control the target device to execute the control strategy corresponding to the target score.

[0032] In this embodiment, the control strategy may include the following: controlling the target device to work in continuous talking mode for a first preset duration, controlling the target device to shut down continuous talking mode, and controlling the target device to work in continuous talking mode for a second preset duration, wherein the second preset duration is shorter than the first preset duration.

[0033] The first preset duration can be set to a slightly longer duration, which means the target device can operate in continuous speaking mode for an extended period. The second preset duration is shorter than the first preset duration, which means the target device can operate in continuous speaking mode for a shorter period.

[0034] This embodiment provides a device control method that, when a target device receives a target voice command, simultaneously acquires the current environmental information of the target device and the corresponding historical speech recognition information of the target device. Simultaneously, it determines the target correlation degree between the target voice command and the historical voice command set received by the target device. By fusing environmental information, historical speech recognition information, and target correlation degree, it obtains a target score corresponding to the target device and, based on this target score, determines a control strategy for controlling the continuous speaking mode of the target device. This achieves multi-dimensional comprehensive management and control of the continuous speaking mode, adapting simultaneously to environmental changes and the voice interaction of the target device. It avoids the inaccuracy of continuous speaking mode control caused by relying solely on single environmental noise, improving the accuracy and reliability of continuous speaking mode control and ensuring a better user experience.

[0035] Referring to Figure 2, which is a flowchart illustrating another device control method provided in an embodiment of this application, the device control method provided in this application includes the following steps: S201: When the target device receives a target voice command, it acquires the environmental information of the current environment of the target device and the historical voice recognition information corresponding to the target device.

[0036] In this embodiment, the environmental information includes the environmental noise level and the number of people in the environment. The method for obtaining the environmental information and historical speech recognition information is the same as in step S101 above, and can be referred to the description in step S101 above. This embodiment will not repeat it here.

[0037] S202: Determine the target historical voice command from the historical voice command set.

[0038] S203: Determine the first semantic correlation between the target speech command and the target historical speech command.

[0039] S204: Determine the second logical correlation between the target voice command and the historical voice command set.

[0040] S205: Based on the first correlation degree and the second correlation degree, determine the target correlation degree between the target voice command and the historical voice command set received by the target device.

[0041] Regarding steps S203 to S205 above, all historical voice commands in the historical voice command set are detected by the target device within the second preset time period and are successfully recognized by the target device. The second preset time period is located before the current moment and is temporally continuous with the current moment. The second preset time period can be set according to actual needs. The main purpose of setting the second preset time period is to ensure that the target relevance determined subsequently depends only on the context of the latest voice interaction, avoiding outdated or obsolete historical voice commands from interfering with the current decision. Successful recognition can be understood as the target device completing an accurate conversion from speech to text for a detected voice command and parsing a clear and executable user intent. The historical voice commands that are successfully recognized by the target device serve to ensure the reliability and effectiveness of the target relevance determined subsequently. The first relevance can be understood as the degree of semantic similarity between the target voice command and the target historical voice commands. The second relevance can be understood as the quantified value of the logical consistency between the target voice command and the entire historical voice command set in terms of the controlled object, the function of the controlled object, and the functional parameters. It can compensate for the problem of some commands having large semantic differences but logical consistency. For example, if the target voice command is to turn up the volume of the living room music, then the controlled object is the living room music, the function of the controlled object is the volume of the living room music, and the function parameter is to turn up the volume.

[0042] After the target device receives the target voice command, based on the current time, it obtains the set of historical voice commands received by the target device within a second preset time period prior to the current time. From the historical voice command set, it identifies the historical voice command corresponding to the longest time interval and designates this historical voice command as the target historical voice command. This target historical voice command is actually the previous voice command successfully recognized by the target device. The texts of the target voice command and the target historical voice command are converted into bag-of-words vectors or sentence embedding vectors to calculate the cosine similarity between the target voice command and the target historical voice command. This cosine similarity is determined as the first correlation between the target voice command and the target historical voice command.

[0043] Based on the control object, the function of the control object, and the function parameters, the logical matching degree between the target speech command and each historical speech command in the historical speech command set is calculated. Then, based on all the obtained logical matching degrees, the second correlation degree between the target speech command and the historical speech command set is calculated.

[0044] After obtaining the first and second correlation scores, they are fused to determine the target correlation score between the target voice command and the historical voice command set received by the target device. Based on this target correlation score, the target score of the target device is then determined. In this embodiment, by calculating the first semantic correlation score between the target voice command and the previous historical voice command successfully recognized by the target device, and calculating the second logical correlation score between the target voice command and the historical voice command set successfully recognized by the target device within a second preset time period, the final target correlation score is determined by fusing the first and second correlation scores. This achieves accurate determination of the target correlation score, avoids interference from long-standing invalid voice commands and biases inherent in single-dimensional judgments, improves the reliability of the determined target correlation score, provides accurate decision-making basis for subsequent continuous speaking mode control, and optimizes the user's voice interaction experience.

[0045] The S204 step described above specifically includes: extracting features from the target speech command and each historical speech command in the historical speech command set based on a preset feature dimension set, to obtain the target features of the target speech command under each preset feature dimension in the preset feature dimension set and the historical features of each historical speech command in the historical speech command set under each preset feature dimension in the preset feature dimension set; determining the logical matching degree between the target speech command and each historical speech command in the historical speech command set based on all target features corresponding to the target speech command and all historical features corresponding to each historical speech command in the historical speech command set; determining the target weight of the logical matching degree corresponding to each historical speech command based on the historical reception time corresponding to each historical speech command in the historical speech command set, wherein the target weight decreases as the time distance between the historical reception time and the current time increases; and using all the obtained target weights, performing a weighted summation of all the obtained logical matching degrees to obtain the second correlation degree between the target speech command and the historical speech command set.

[0046] In the above, the preset feature dimension set consists of the control object, the function of the control object, and the function parameters. The control object, the function of the control object, and the function parameters are consistent with those described above, and can be found in the above description. Historical reception time can be understood as the time when the target device historically received historical voice commands.

[0047] For each historical speech command in the target speech command and historical speech command sets, regularized feature extraction is used to extract the target features of the target speech command under each preset feature dimension in the preset feature dimension set, and regularized feature extraction is used to extract the historical features of the historical speech command under each preset feature dimension in the preset feature dimension set.

[0048] After obtaining all target features corresponding to the target speech command and all historical features corresponding to each historical speech command, for each historical speech command, all overlapping features between all historical features corresponding to the historical speech command and all target features of the target speech command are determined, and the priority of the preset feature dimension to which the overlapping features belong is determined. Based on each priority, the logical matching score of the corresponding overlapping feature is determined. The logical matching degree between the target speech command and the historical speech command is obtained by adding all the logical matching scores. Among them, the control object in the preset feature dimension set has the highest priority, followed by the function of the control object, and the function parameter has the lowest priority. Each preset feature dimension has a corresponding logical matching score pre-set based on the priority.

[0049] In the above example, all target features corresponding to the target voice command include target feature A1 corresponding to the control object, target feature A2 corresponding to the function of the control object, and target feature A3 corresponding to the function parameters. All historical features corresponding to the historical voice command include historical feature B1 corresponding to the control object, target feature B2 corresponding to the function of the control object, and target feature B3 corresponding to the function parameters. If target feature A1 and historical feature B1 are overlapping features, then the overlapping feature belongs to the preset feature dimension of the control object. Therefore, the pre-set logical matching score corresponding to the control object is determined as the logical matching degree between the target voice command and the historical voice command.

[0050] After obtaining the logical matching scores, a corresponding target weight is assigned to the logical matching score corresponding to each historical voice command according to the time decay rule. The target weight decreases as the time distance between the historical reception time corresponding to the historical voice command and the current time increases. In other words, the closer the historical reception time corresponding to the historical voice command is to the current time, the higher the target weight. The time decay rule can be set according to actual needs, and will not be elaborated here in this embodiment.

[0051] After obtaining the target weights corresponding to each logical matching degree, each logical matching degree is multiplied by its corresponding target weight to obtain a product. The sum of all products yields the second correlation degree between the target voice command and the historical voice command set. This embodiment provides a method for determining the second correlation degree by using a preset feature dimension set. Features of the target voice command and each historical voice command under each preset dimension of the preset feature dimension set are extracted. The logical matching degree between the target voice command and the historical voice commands is calculated line by line, and a weighted sum is obtained by combining the target weights that decay over time to achieve the calculation of the second correlation degree. This allows for a more comprehensive evaluation of the logical consistency between the target voice command and the historical voice command set, improving the accuracy of the determined second correlation degree. This provides a reliable basis for subsequent target correlation degree calculation and continuous speaking mode control, further enhancing the user experience.

[0052] In the above, step S205 specifically includes: determining a first comparison result between a first correlation degree and a first correlation degree threshold, and a second comparison result between a second correlation degree and a second correlation degree threshold; based on the first comparison result and the second comparison result, determining a fourth weight corresponding to the first correlation degree and a fifth weight corresponding to the second correlation degree; using the fourth weight and the fifth weight, performing a weighted summation on the first correlation degree and the second correlation degree to obtain the target correlation degree between the target voice command and the historical voice command set received by the target device.

[0053] The first relevance threshold is used to classify the semantic relevance strength. This threshold can be set according to actual needs. For example, when the first relevance threshold is 0.6, a relevance score greater than or equal to 0.6 indicates a high semantic relevance, while a relevance score less than 0.6 indicates a low semantic relevance. The second relevance threshold is used to classify the logical relevance strength. This threshold can also be set according to actual needs. For example, when the second relevance threshold is 0.7, a relevance score greater than or equal to 0.7 indicates a high logical relevance, while a relevance score less than 0.7 indicates a low logical relevance.

[0054] After obtaining the first and second correlation scores, a first comparison result between the first correlation score and a first correlation score threshold, and a second comparison result between the second correlation score and a second correlation score threshold are determined. Based on the pre-set mapping relationship between the first and second comparison results and the weights corresponding to the first and second correlation scores, a fourth weight corresponding to the first correlation score and a fifth weight corresponding to the second correlation score are determined under the first and second comparison results. The product between the first correlation score and the fourth weight and the product between the second correlation score and the fifth weight are determined, and the sum of the two products is determined as the target correlation score between the target voice command and the historical voice command set detected by the target device. In this way, this embodiment dynamically allocates the weights of the first and second correlation scores through dual-threshold comparison, integrates dual-dimensional calculation of the target correlation score, avoids misjudgment of the target correlation score caused by fixed weights, improves the accuracy of the determined target correlation score, provides a reliable basis for continuous speaking mode control, and further improves the user experience.

[0055] S206: Determine the first weight corresponding to environmental information, the second weight corresponding to historical speech recognition information, and the third weight corresponding to target correlation.

[0056] S207: Using the first weight, the second weight, and the third weight, the environmental information, historical speech recognition information, and target relevance are weighted and summed to obtain the target score corresponding to the target device.

[0057] Regarding steps S206 and S207 above, the first weight is a quantized proportional coefficient assigned to environmental information when determining the target score corresponding to the target device, measuring the degree of influence of environmental information on the continuous speech mode control decision. The second weight is a quantized proportional coefficient assigned to historical speech recognition information when determining the target score corresponding to the target device, measuring the degree of influence of historical speech recognition information on the continuous speech mode control decision. The third weight is a quantized proportional coefficient assigned to target relevance when determining the target score corresponding to the target device, measuring the degree of influence of target relevance on the continuous speech mode control decision. The sum of the first, second, and third weights is 1.

[0058] Specifically, when determining the first, second, and third weights, a pre-set mapping table between environmental information, historical speech recognition information, target relevance, and weights can be used to determine the first weight corresponding to the environmental information, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target relevance. Alternatively, the first weight corresponding to the environmental information, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target relevance can be dynamically determined based on the target device's historical voice interaction information. Specific implementation methods are described below.

[0059] After obtaining the first, second, and third weights, the environmental information, historical speech recognition information, and target relevance are quantified to obtain the environmental score corresponding to the environmental information, the speech recognition score corresponding to the historical speech recognition information, and the relevance score corresponding to the target relevance. The first product between the first weight and the environmental score, the second product between the second weight and the speech recognition score, and the third product between the third weight and the relevance score are determined. The sum of these three products is then used as the weighted sum to obtain the target score corresponding to the target device. It should be noted that when the environmental information includes the environmental noise level and the number of people in the environment, the environmental score decreases as the environmental noise level increases, and the environmental score decreases as the number of people in the environment increases. The speech recognition score increases as the success rate of speech recognition increases, or decreases as the failure rate of speech recognition increases. The relevance score increases as the target relevance increases. Through the above methods, this embodiment provides a way to determine the target score. By configuring corresponding weights for environmental information, historical speech recognition information, and target relevance, a weighted summation is achieved to realize the orderly fusion of multi-dimensional information, thereby providing a basis for the precise control of the subsequent continuous speaking mode, thus improving the scene adaptability of the continuous speaking mode and optimizing the user's voice interaction experience.

[0060] In this embodiment, the first weight includes a first sub-weight corresponding to the environmental noise level and a second sub-weight corresponding to the number of environmental objects.

[0061] The above-mentioned step S206 specifically includes: obtaining the initial weights corresponding to the environmental noise level, the number of people in the environment, historical speech recognition information, and target relevance; determining the first number of times the target device has erroneously triggered a speech event within a first preset time period and the second number of times the target device has successfully recognized speech commands consecutively before the current moment; and correcting each initial weight based on the first and second counts to obtain the first sub-weight corresponding to the environmental noise level, the second sub-weight corresponding to the number of people in the environment, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target relevance.

[0062] The first sub-weight is a quantized proportional coefficient assigned to the environmental noise level when determining the target score corresponding to the target device, measuring the degree of influence of the environmental noise level on the continuous speech mode control decision. The second sub-weight is a quantized proportional coefficient assigned to the number of environmental objects when determining the target score corresponding to the target device, measuring the degree of influence of the number of environmental objects on the continuous speech mode control decision. The sum of the first sub-weight, the second sub-weight, the second weight, and the third weight is 1. The initial weights include the initial weights preset for the environmental noise level, the initial weights preset for the number of environmental objects, the initial weights preset for historical speech recognition information, and the initial weights preset for the target correlation. The sum of all the initial weights is also 1. The first preset time period is before the current time and is temporally continuous with the current time. The first preset time period can be set according to actual needs, and this embodiment does not make specific limitations here. A false triggering of a voice event by the target device can be understood as an event in which the target device is woken up to perform speech recognition without the user's intention to interact.

[0063] After obtaining the initial weights corresponding to the environmental noise level, the number of people in the environment, historical speech recognition information, and target relevance, the number of times the target device erroneously triggers a speech event within the first preset time period is counted to obtain the first count, and the number of times the target device has successfully recognized speech commands consecutively before the current moment is counted to obtain the second count.

[0064] After obtaining the first and second counts, the target frequency is compared with the threshold of the first count to obtain the third comparison result. The second count is compared with the threshold of the second count to obtain the fourth comparison result. Using the pre-set mapping rules between the third and fourth comparison results and the correction values ​​of each initial weight, the correction values ​​of each initial weight corresponding to the third and fourth comparison results are determined. Then, each correction value is used to correct the corresponding initial weight to obtain the first sub-weight, the second sub-weight, the second weight, and the third weight.

[0065] It should be noted that the above mapping rules must follow these principles: When the first number of occurrences exceeds the threshold for the first occurrence, it indicates a harsh environment. In this case, the initial weights of environmental noise level and the number of objects in the environment should be increased, while the initial weights of historical speech recognition information and target relevance should be decreased. When the second number of occurrences exceeds the threshold for the second occurrence, it indicates smooth speech interaction. In this case, the initial weights of historical speech recognition information and target relevance should be increased, while the initial weights of environmental noise level and the number of objects in the environment should be decreased. This ensures that in harsh environments, increasing the environmental weights amplifies the negative impact of the harsh environment score on the target score, thus lowering the target score. Conversely, in smooth speech interaction, increasing the weights of historical speech recognition information and target relevance amplifies the positive impact of smooth speech interaction on the target score, thus raising the target score.

[0066] In this embodiment, when determining each weight, the first number of times the target device erroneously triggers a voice event within a first preset time period and the second number of times the target device has successfully recognized voice commands consecutively before the current moment are used to adaptively and dynamically determine each weight. This ensures that the determined weights match the current environmental state and the voice interaction of the target device, thereby guaranteeing the accuracy of the subsequently determined target score and providing a more reliable basis for the control of the continuous speaking mode.

[0067] S208: Control the target device to execute the control strategy corresponding to the target score.

[0068] In this embodiment, step S208 specifically includes: when the target score is greater than or equal to a first score threshold, controlling the target device to operate in continuous speaking mode for a first preset duration; when the target score is less than a second score threshold, controlling the continuous speaking mode of the target device to be turned off, the second score threshold being less than the first score threshold; when the target score is greater than or equal to the second score threshold and less than the first score threshold, controlling the target device to operate in continuous speaking mode for a second preset duration, the second preset duration being less than the first preset duration.

[0069] The first score threshold is a pre-set, relatively high score threshold, serving as the decision boundary for the target device to operate in continuous speaking mode for an extended period. The second score threshold is lower than the first score threshold, serving as the decision boundary for disabling the target device's continuous speaking mode. The first preset duration threshold is a pre-set, relatively long duration for maintaining continuous speaking mode, providing users with a seamless window for continuous interaction. The second preset duration threshold is a pre-set, relatively short duration for maintaining continuous speaking mode, providing users with a temporary window for continuous interaction to prevent the target device from experiencing a degraded user experience or triggering accidental actions due to prolonged waiting for invalid input.

[0070] After obtaining the target score, it is compared with a first score threshold and a second score threshold. If the target score is greater than or equal to the first score threshold, the target device is controlled to operate in continuous speaking mode for a first duration to maximize the fluency of voice interaction. If the target score is less than the second score threshold, the continuous speaking mode of the target device is turned off. It should be noted that when the target score is less than the second score threshold, a fault-tolerant prompt is triggered to improve the user experience, for example, "I didn't quite hear you, could you say it again?" After triggering the fault-tolerant prompt, the target device continues to operate in continuous speaking mode, waiting to receive the next target voice command. After receiving the next target voice command, step S201 is executed. If the number of fault-tolerant prompts triggered exceeds a preset threshold, the continuous speaking mode of the target device is turned off. If the target score is equal to or equal to the second score threshold and less than the first score threshold, the target device is controlled to operate in continuous speaking mode for a second duration to still receive user voice commands within a shortened time window.

[0071] It should be noted that, in order to improve the accuracy of the continuous speaking mode of the target device, the third number of false voice events triggered by the target device within the fourth preset time period is determined. When the third number is greater than the third number threshold, the obtained first and second score thresholds are increased by a preset step size to update the first and second score thresholds. The updated first and second score thresholds are then used to execute the control strategy steps corresponding to the target score. When the third number is less than or equal to the third number threshold, the obtained first and second score thresholds are used to execute the control strategy steps corresponding to the target score. The fourth preset time period is located before the current time and is temporally continuous with the current time. False trigger events can be referred to as described above, and will not be repeated here in this embodiment.

[0072] Through the above methods, this embodiment compares the target score with different score thresholds to achieve precise control of the continuous speaking mode based on the comparison results. This allows the control results of the continuous speaking mode to adapt to the current environment and the user's voice interaction needs, thus ensuring the user experience.

[0073] This embodiment provides a device control method that, when a target device receives a target voice command, simultaneously acquires the current environmental information of the target device and the corresponding historical speech recognition information of the target device. Simultaneously, it determines the target correlation degree between the target voice command and the historical voice command set received by the target device. By fusing environmental information, historical speech recognition information, and target correlation degree, it obtains a target score corresponding to the target device and, based on this target score, determines a control strategy for controlling the continuous speaking mode of the target device. This achieves multi-dimensional comprehensive management and control of the continuous speaking mode, adapting simultaneously to environmental changes and the voice interaction of the target device. It avoids the inaccuracy of continuous speaking mode control caused by relying solely on single environmental noise, improving the accuracy and reliability of continuous speaking mode control and ensuring a better user experience.

[0074] Referring to Figure 3, which is a schematic diagram of the structure of a device control apparatus provided in an embodiment of this application, the device control apparatus provided in this application includes an acquisition module 10, a determination module 20, and a control module 30. The acquisition module 10 is used to acquire environmental information of the current environment of the target device and historical speech recognition information corresponding to the target device when the target device receives a target voice command. The determination module 20 is used to determine the target correlation degree between the target voice command and the set of historical voice commands received by the target device based on the target voice command. The determination module 20 is also used to determine a target score corresponding to the target device based on the environmental information, the historical speech recognition information, and the target correlation degree. The target score is used to indicate a control strategy for controlling the continuous speaking mode of the target device. The control module 30 is used to control the target device to execute the control strategy corresponding to the target score.

[0075] In this embodiment, the determining module 20 is further configured to: determine the first weight corresponding to the environmental information, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target correlation; and use the first weight, the second weight, and the third weight to perform a weighted summation on the environmental information, the historical speech recognition information, and the target correlation to obtain the target score corresponding to the target device.

[0076] In this embodiment, the environmental information includes the environmental noise level and the number of environmental objects. The first weight includes a first sub-weight corresponding to the environmental noise level and a second sub-weight corresponding to the number of environmental objects. The determining module 20 is further configured to: obtain the initial weights corresponding to the environmental noise level, the number of environmental objects, the historical speech recognition information, and the target correlation; determine the first number of times the target device has erroneously triggered a speech event within a first preset time period and the second number of times the target device has successfully recognized a speech command consecutively before the current moment; and correct each of the initial weights according to the first number and the second number to obtain the first sub-weight corresponding to the environmental noise level, the second sub-weight corresponding to the number of environmental objects, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target correlation.

[0077] In this embodiment, all historical voice commands in the historical voice command set are received by the target device within a second preset time period and are all successfully recognized by the target device; the determining module 20 is further configured to: determine a target historical voice command from the historical voice command set, wherein the target historical voice command is the previous historical voice command of the target voice command; determine a first semantic correlation between the target voice command and the target historical voice command; determine a second logical correlation between the target voice command and the historical voice command set; and, based on the first correlation and the second correlation, determine a target correlation between the target voice command and the historical voice command set received by the target device.

[0078] In this embodiment, the determining module 20 is further configured to: extract features from the target voice command and each of the historical voice commands in the historical voice command set based on a preset feature dimension set, to obtain target features of the target voice command under each preset feature dimension in the preset feature dimension set and historical features of each of the historical voice commands in the historical voice command set under each preset feature dimension in the preset feature dimension set; determine the logical matching degree between the target voice command and each of the historical voice commands in the historical voice command set based on all the target features corresponding to the target voice command and all the historical features corresponding to each of the historical voice commands in the historical voice command set; determine the target weight of the logical matching degree corresponding to each of the historical voice commands based on the historical reception time corresponding to each of the historical voice commands in the historical voice command set, wherein the target weight decreases as the time distance between the historical reception time and the current time increases; and perform a weighted summation of all the obtained logical matching degrees using all the obtained target weights to obtain a second correlation degree between the target voice command and the historical voice command set.

[0079] In this embodiment, the determining module 20 is further configured to: determine a first comparison result between the first correlation degree and the first correlation degree threshold, and a second comparison result between the second correlation degree and the second correlation degree threshold; determine a fourth weight corresponding to the first correlation degree and a fifth weight corresponding to the second correlation degree based on the first comparison result and the second comparison result; and perform a weighted summation of the first correlation degree and the second correlation degree using the fourth weight and the fifth weight to obtain the target correlation degree between the target voice command and the historical voice command set received by the target device.

[0080] In this embodiment, the control module 30 is further configured to: control the target device to operate in the continuous speaking mode for a first preset duration when the target score is greater than or equal to a first score threshold; control the continuous speaking mode of the target device to be turned off when the target score is less than a second score threshold, wherein the second score threshold is less than the first score threshold; and control the target device to operate in the continuous speaking mode for a second preset duration when the target score is greater than or equal to the second score threshold and less than the first score threshold, wherein the second preset duration is less than the first preset duration.

[0081] This embodiment provides a device control apparatus that, when a target device receives a target voice command, simultaneously acquires the current environmental information of the target device and the corresponding historical speech recognition information of the target device. It also determines the target correlation degree between the target voice command and the historical voice command set received by the target device. By fusing environmental information, historical speech recognition information, and target correlation degree, it obtains a target score corresponding to the target device and controls the continuous speaking mode of the target device based on the control strategy determined by the target score. This achieves multi-dimensional comprehensive management and control of the continuous speaking mode, adapting to both environmental changes and the voice interaction of the target device. It avoids the problem of inaccurate control of the continuous speaking mode caused by relying solely on single environmental noise, improving the accuracy and reliability of the continuous speaking mode control and ensuring a better user experience.

[0082] Figure 4 is a schematic diagram of a smart home device provided in an embodiment of this application. The smart home device 400 shown in Figure 4 includes: at least one processor 401, a memory 402, at least one network interface 404, and other user interfaces 403. The various components in the smart home device 400 are coupled together through a bus system 405. It is understood that the bus system 405 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 405 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 405 in Figure 4.

[0083] The user interface 403 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).

[0084] It is understood that the memory 402 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 402 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0085] In some implementations, memory 402 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 4021 and application program 4022.

[0086] The operating system 4021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 4022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 4022.

[0087] In this embodiment of the invention, the processor 401 executes the method steps provided in each method embodiment by calling the program or instructions stored in the memory 402, specifically the program or instructions stored in the application program 4022.

[0088] The methods disclosed in the above embodiments of the present invention can be applied to processor 401, or implemented by processor 401. Processor 401 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 401 or by instructions in the form of software. The processor 401 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 402. Processor 401 reads the information in memory 402 and, in conjunction with its hardware, completes the steps of the above method.

[0089] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0090] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0091] The smart home device provided in this embodiment can be the smart home device shown in Figure 4. It can execute all the steps of the control method of the device shown in Figures 1 and 2, thereby achieving the technical effect of the control method of the device shown in Figures 1 and 2. For details, please refer to the relevant descriptions in Figures 1 and 2. For the sake of brevity, it will not be elaborated here.

[0092] This invention also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.

[0093] When one or more programs in the storage medium can be executed by one or more processors to implement the device control method described above, which is executed on the device control device side.

[0094] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0095] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0096] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for controlling a device, characterized in that, include: When the target device receives a target voice command, it acquires environmental information about the current environment of the target device and historical speech recognition information corresponding to the target device; based on the target voice command, it determines the target correlation degree between the target voice command and the set of historical voice commands received by the target device; based on the environmental information, the historical speech recognition information, and the target correlation degree, it determines the target score corresponding to the target device, the target score being used to indicate the control strategy for controlling the continuous speaking mode of the target device; and it controls the target device to execute the control strategy corresponding to the target score.

2. The method according to claim 1, characterized in that, The step of determining the target score corresponding to the target device based on the environmental information, the historical speech recognition information, and the target correlation degree includes: determining a first weight corresponding to the environmental information, a second weight corresponding to the historical speech recognition information, and a third weight corresponding to the target correlation degree; and using the first weight, the second weight, and the third weight, performing a weighted summation on the environmental information, the historical speech recognition information, and the target correlation degree to obtain the target score corresponding to the target device.

3. The method according to claim 2, characterized in that, The environmental information includes environmental noise level and number of environmental objects. The first weight includes a first sub-weight corresponding to the environmental noise level and a second sub-weight corresponding to the number of environmental objects. Determining the first weight corresponding to the environmental information, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target correlation includes: obtaining the initial weights corresponding to the environmental noise level, the number of environmental objects, the historical speech recognition information, and the target correlation; determining the first number of times the target device has erroneously triggered a voice event within a first preset time period and the second number of times the target device has successfully recognized voice commands consecutively before the current moment; and correcting each of the initial weights based on the first number and the second number to obtain the first sub-weight corresponding to the environmental noise level, the second sub-weight corresponding to the number of environmental objects, the second weight corresponding to the historical speech recognition information, and the third weight corresponding to the target correlation.

4. The method according to claim 1, characterized in that, All historical voice commands in the historical voice command set are received by the target device within a second preset time period and are all successfully recognized by the target device. Determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the target voice command includes: identifying a target historical voice command from the historical voice command set, where the target historical voice command is the previous historical voice command of the target voice command; determining a first semantic correlation degree between the target voice command and the target historical voice command; determining a second logical correlation degree between the target voice command and the historical voice command set; and determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the first correlation degree and the second correlation degree.

5. The method according to claim 4, characterized in that, The step of determining the second logical correlation between the target voice command and the historical voice command set includes: extracting features from the target voice command and each historical voice command in the historical voice command set based on a preset feature dimension set to obtain target features of the target voice command under each preset feature dimension in the preset feature dimension set and historical features of each historical voice command in the historical voice command set under each preset feature dimension in the preset feature dimension set; determining the logical matching degree between the target voice command and each historical voice command in the historical voice command set based on all target features corresponding to the target voice command and all historical features corresponding to each historical voice command in the historical voice command set; determining the target weight of the logical matching degree corresponding to each historical voice command based on the historical reception time corresponding to each historical voice command in the historical voice command set, wherein the target weight decreases as the time distance between the historical reception time and the current time increases; and using all the obtained target weights, performing a weighted summation of all the obtained logical matching degrees to obtain the second logical correlation between the target voice command and the historical voice command set.

6. The method according to claim 4, characterized in that, Determining the target correlation degree between the target voice command and the historical voice command set received by the target device based on the first correlation degree and the second correlation degree includes: determining a first comparison result between the first correlation degree and a first correlation degree threshold, and a second comparison result between the second correlation degree and a second correlation degree threshold; determining a fourth weight corresponding to the first correlation degree and a fifth weight corresponding to the second correlation degree based on the first comparison result and the second comparison result; and using the fourth weight and the fifth weight to perform a weighted summation of the first correlation degree and the second correlation degree to obtain the target correlation degree between the target voice command and the historical voice command set received by the target device.

7. The method according to claim 1, characterized in that, The control strategy corresponding to the target score for controlling the target device includes: when the target score is greater than or equal to a first score threshold, controlling the target device to operate in the continuous speaking mode for a first preset duration; when the target score is less than a second score threshold, controlling the continuous speaking mode of the target device to be turned off, wherein the second score threshold is less than the first score threshold; and when the target score is greater than or equal to the second score threshold and less than the first score threshold, controlling the target device to operate in the continuous speaking mode for a second preset duration, wherein the second preset duration is less than the first preset duration.

8. A control device for an equipment, characterized in that, include: The acquisition module is used to acquire environmental information of the current environment of the target device and historical speech recognition information corresponding to the target device when the target device receives the target voice command; The determining module is configured to determine the target correlation degree between the target voice command and the historical voice command set received by the target device based on the target voice command; the determining module is further configured to determine the target score corresponding to the target device based on the environmental information, the historical voice recognition information and the target correlation degree, the target score being used to indicate the control strategy for controlling the continuous speaking mode of the target device; the control module is configured to control the target device to execute the control strategy corresponding to the target score.

9. A smart home device, comprising: A processor and a memory, the processor being configured to execute a device control program stored in the memory to implement the device control method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the control method of the device according to any one of claims 1 to 7.