Smart home voice control method and system under heterogeneous resource condition

By using preset voice recognition models and appropriate communication methods in the smart home voice control system, heterogeneous resources are integrated to realize whole-house voice control, the problem of difficulty in taking into account existing home devices in the prior art is solved, and user experience and compatibility are improved.

CN120126474APending Publication Date: 2025-06-10GUANGDONG ELECTRIC POWER SCI RES INST ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318358.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

It is difficult for smart home voice control systems to take into account existing home devices, resulting in low recognition rate and poor user experience.

Method used

The user's voice commands are identified through the preset voice recognition model, the target device and operation commands are determined, and the appropriate communication method is selected according to the device type for control, integrating heterogeneous resources to realize whole-house voice control.

Benefits of technology

Improves the compatibility and user experience of smart home voice control, ensuring effective voice control for various home devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126474A_ABST
    Figure CN120126474A_ABST
Patent Text Reader

Abstract

The invention provides a smart home voice control method and system under a heterogeneous resource condition. The method comprises the following steps: acquiring a voice instruction of a user; inputting the voice instruction into a preset voice recognition model, so that the voice recognition model recognizes the content of the voice instruction, and further determining target equipment to be controlled by a user and a corresponding operation instruction; according to the type of the target device, a communication mode with the target device is determined, and the type of the target device comprises an intelligent home device, a home device with a Bluetooth communication or carrier communication function, a home device with a voice recognition function and a traditional home device; and controlling the target equipment to execute the operation instruction based on the communication mode, and realizing voice control of the whole house home equipment by using different communication control modes, thereby improving the compatibility of smart home voice control and the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of voice control and smart home, and particularly to a smart home voice control method and system under heterogeneous resource conditions. Background Art

[0002] A smart home system is a system that integrates relevant facilities in home life into a unified management system by using advanced technologies such as comprehensive wiring and network communication. It can realize various functions such as lighting control, curtain control, air-conditioning control, air quality detection, security linkage, and remote control, providing an intelligent, convenient, and flexible usage experience for home life.

[0003] With the continuous development of artificial intelligence technology, smart home control systems have gradually become an indispensable part of people's lives. In this field, intelligent voice recognition technology plays an increasingly important role. Intelligent voice recognition technology is a technology that performs corresponding operations by recognizing and understanding human voice information. In a smart home control system, this technology allows users to control various home devices, such as lighting, air-conditioning, and television, through voice commands. Intelligent voice recognition technology includes key links such as voice signal acquisition, feature extraction, and model training. In a smart home control system, the main application scenarios of intelligent voice recognition technology include smart speakers and smart home control panels. A smart speaker can receive a user's voice command, convert it into a corresponding control signal, and then send it to a home device. A smart home control panel can match a user's voice command with a control signal of a home device to achieve more precise control.

[0004] Intelligent voice recognition technology has obvious advantages in a smart home control system. First of all, it can achieve various complex operations through voice commands, improving the convenience of user use. Secondly, intelligent voice recognition technology can reduce the probability of misoperation and improve the accuracy of control. In addition, this technology can also support multiple languages and dialects to meet the needs of different users. However, intelligent voice recognition technology also faces some challenges in a smart home control system. First of all, the voice recognition rate is one of the key factors affecting the user experience. Although intelligent voice recognition technology is constantly developing, it may still be affected by factors such as noise, accent, and speech rate, resulting in the recognition rate not reaching 100%. And a smart home system not only needs to achieve intelligent control with brand-new high-end complete sets of equipment in newly added rooms; at the same time, it should also take into account existing home devices and be able to achieve full-house intelligence at the lowest cost. Summary of the Invention

[0005] In view of the above technical problems, the present application provides a smart home voice control method and system under heterogeneous resource conditions, which realizes voice control of home appliances throughout the house by using different communication control methods, improving the compatibility of smart home voice control and the user experience.

[0006] In a first aspect, an embodiment of the present application provides a smart home voice control method under heterogeneous resource conditions, including:

[0007] Obtain a voice command from the user;

[0008] Input the voice command into a preset voice recognition model, so that the voice recognition model recognizes the content of the voice command, and then determines the target device to be controlled by the user and the corresponding operation command;

[0009] According to the type of the target device, determine the communication method with the target device, where the type of the target device includes intelligent home appliances, home appliances with Bluetooth communication or carrier communication functions, home appliances with voice recognition functions, and traditional home appliances;

[0010] Based on the communication method, control the target device to execute the operation command.

[0011] An embodiment of the present application provides a smart home voice control method under heterogeneous resource conditions. By using a preset voice recognition model to recognize the voice command of the user, the target device to be operated by the user and the operation command to be executed by the target device are determined. Then, according to the type of the target device, the communication method with the target device is determined, and further the target device is controlled to execute the operation command. Through the use of different communication control methods, the heterogeneous resources throughout the house are integrated in the embodiment of the present application. Whether it is intelligent home appliances, home appliances with Bluetooth communication or carrier communication functions, home appliances with voice recognition functions, or traditional home appliances, voice control can be realized through the embodiment of the present application, solving the technical problem that the smart home voice control in the prior art cannot take into account the existing home appliances, and improving the compatibility of smart home voice control and the user experience.

[0012] In a possible implementation manner, the voice recognition model recognizes the content of the voice command, including:

[0013] The voice recognition model generates a first recognition result according to the voice command;

[0014] Input the first recognition result into a preset recognition result evaluation model, so that the recognition result evaluation model scores the confidence of the first recognition result and generates a corresponding confidence score value;

[0015] If the confidence score value is greater than or equal to the confidence threshold, then use the first recognition result as the content of the voice command;

[0016] If the confidence score value is less than the confidence threshold, then upload the voice command to the cloud so that the cloud recognition model pre-deployed on the cloud generates a corresponding second recognition result according to the voice command;

[0017] Obtain the second recognition result and use the second recognition result as the content of the voice command.

[0018] In an embodiment of the present application, a voice recognition method is provided. By performing a confidence score on the first recognition result of a voice recognition model, the accuracy of the first recognition result is evaluated. When the confidence score value is low, the prior art usually responds with a voice broadcast to inform the user to reissue the voice command. However, this method cannot guarantee that the user's next voice command can be accurately recognized, which may lead to the user repeatedly issuing the same voice command. Therefore, the embodiment of the present application improves the accuracy of voice recognition through a cloud-edge collaboration method, and deploys a more complex and more accurate cloud recognition model on the cloud to recognize the user's voice command. At the same time, the embodiment of the present application takes into account both the recognition response speed and the recognition accuracy. By comparing the confidence score value with the confidence threshold, it is determined whether to start the cloud recognition model for secondary recognition. When the confidence is high, the recognition result of the voice recognition model is directly output. When the confidence is low, secondary recognition is performed through the cloud recognition model, giving full play to the advantages of the high efficiency of the local voice recognition model and the accuracy of the cloud recognition model, and improving the user experience.

[0019] Further, the confidence threshold is initialized based on the historical data of other voice recognition models and updated according to the historical recognition data of the current voice recognition model, including:

[0020] Generate an initial confidence threshold according to the historical data of other voice recognition models and set it in the current voice recognition model;

[0021] After a preset update period, obtain the historical recognition data of the current voice recognition model during the update period;

[0022] Analyze the historical recognition data, and count the number of recognition errors of the current voice recognition model and the number of times the inference time of the cloud recognition model exceeds the preset time;

[0023] Generate corresponding first evaluation indicators and second evaluation indicators according to the number of recognition errors of the voice recognition model and the number of times the inference time of the cloud recognition model exceeds the preset time;

[0024] Update the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold, and set the confidence threshold in the current speech recognition model;

[0025] The confidence threshold is updated regularly according to the update period.

[0026] Further, updating the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold includes:

[0027] Calculate a comprehensive evaluation index according to the preset index weight, the first evaluation index and the second evaluation index;

[0028] Increase or decrease the initial confidence threshold according to the comprehensive evaluation index to obtain the confidence threshold.

[0029] The embodiment of the present application provides a method for setting and updating the confidence threshold. When the speech recognition model is first deployed, the current speech recognition model has no historical data as a reference. Therefore, an initial confidence threshold is generated according to the historical data of other speech recognition models and set in the current speech recognition model. Since the usage environments and target users of each speech recognition model are different, the initial confidence threshold needs to be updated according to the actual usage situation after being used for a period of time to improve the recognition efficiency and accuracy of the current speech recognition model. Therefore, after the current speech recognition model runs for a period of time, a first evaluation index and a second evaluation index are generated according to the historical recognition data of the current speech recognition model, and then the initial confidence threshold is updated according to the first evaluation index and the second evaluation index. Among them, the first evaluation index is used to evaluate the recognition accuracy of the model, and the first evaluation index is used to evaluate the recognition efficiency of the model. By increasing or decreasing the initial confidence threshold, the updated confidence threshold is more suitable for the current target user, improving the user experience.

[0030] In a possible implementation manner, when the type of the target device is a home device with Bluetooth communication or carrier communication function, controlling the target device to execute the operation instruction based on the communication method includes:

[0031] Convert the operation instruction into a corresponding control code according to a preset control code correspondence table, and the control code correspondence table is constructed and generated according to the executable operation list of the target device. Each executable operation of the target device corresponds to a digital ID with a preset length, and the digital ID is the control code.

[0032] Generate a corresponding control message according to a preset communication protocol and the control code;

[0033] Send the control message to the target device through Bluetooth communication or carrier communication, so that the target device executes the operation instruction.

[0034] An embodiment of the present application provides a communication control method for a home device with Bluetooth communication or carrier communication functions. By using a preset control code correspondence table, the recognized operation instruction is converted into an operation code recognizable by the target device, and then a control message for communicating with the target device is generated, finally converting the user's voice instruction into a corresponding control message, so that devices without voice recognition can also be controlled by voice instructions, improving the compatibility of smart home voice control and the user experience.

[0035] In a possible implementation manner, when the type of the target device is a home device with voice recognition function, controlling the target device to execute the operation instruction based on the communication method includes:

[0036] Determine the operation that the target device currently needs to execute according to the operation instruction and the list of executable operations of the target device;

[0037] Generate a corresponding AI voice instruction according to the operation that the target device currently needs to execute;

[0038] Broadcast the AI voice instruction to the target device, so that the target device recognizes the AI voice instruction according to the built-in device voice recognition model, and then executes the operation instruction, where the device voice recognition model is trained based on a number of AI voice instructions.

[0039] An embodiment of the present application provides a communication control method for a home device with voice recognition function. Although these home devices have voice recognition function, their computing power and memory are limited. Even if model compression and acceleration methods such as network pruning, parameter quantization, and knowledge distillation are used, it is difficult to support the voice recognition of different users' pronunciation habits, accents, and dialects. And the control center can usually deploy more complex models, with higher voice recognition accuracy than the home device. Therefore, an embodiment of the present application improves the voice recognition accuracy of the home device by means of voice conversion, uses the voice recognition model of the control center to recognize the user's voice instruction instead of these home devices, and then converts it into a corresponding AI voice instruction according to the recognition result. At the same time, the device voice recognition model of the home device is trained in advance with a number of AI voice instructions, so that the home can accurately recognize the AI voice instruction during actual application, improving the voice recognition accuracy of the home device with voice recognition function, and thus improving the user experience.

[0040] Further, the voice recognition model is updated regularly based on a preset time interval, including:

[0041] Collect the historical recognition data of the speech recognition model within the preset time interval, where the historical recognition data includes the user's voice commands, the recognition results of the speech recognition model, and the number of times the user repeats the pronunciation;

[0042] Generate a number of training samples after manually annotating the historical recognition data;

[0043] Upload each of the training samples to the cloud so that the cloud can train and optimize the speech recognition model to be updated in the cloud according to each of the training samples, and generate a corresponding updated speech recognition model, where the model parameters of the speech recognition model to be updated are the same as those of the speech recognition model;

[0044] Obtain the updated speech recognition model sent from the cloud, and replace the speech recognition model with the updated speech recognition model for redeployment.

[0045] The most widely used scenario of the smart home system is the home, and the users are relatively fixed. Their pronunciation habits and accents are also relatively fixed. The initially deployed speech recognition model is a model with strong generalization ability, but it is not fine-tuned for a specific home environment. Therefore, after the smart home system is used for a period of time, the speech recognition model can be updated and optimized according to the collected data and user feedback. In the embodiments of the present application, the periodic update of the speech recognition model is realized by uploading the historical recognition data to the cloud, so that the local speech recognition model is gradually optimized, the accuracy of smart home voice control is improved, and the user experience is further improved.

[0046] In a second aspect, an embodiment of the present application provides a smart home voice control system under heterogeneous resource conditions, including an acquisition module, a speech recognition module, a communication method determination module, and a control module;

[0047] The acquisition module is used to acquire the user's voice command;

[0048] The speech recognition module is used to input the voice command into a preset speech recognition model, so that the speech recognition model recognizes the content of the voice command, and further determines the target device to be controlled by the user and the corresponding operation command;

[0049] The communication method determination module is used to determine the communication method with the target device according to the type of the target device, where the type of the target device includes intelligent home devices, home devices with Bluetooth communication or carrier communication functions, home devices with speech recognition functions, and traditional home devices;

[0050] The control module is used to control the target device to execute the operation command based on the communication method.

[0051] In a possible implementation manner, the speech recognition model recognizes the content of the speech instruction, including:

[0052] The speech recognition model generates a first recognition result according to the speech instruction;

[0053] Input the first recognition result into a preset recognition result evaluation model, so that the recognition result evaluation model scores the confidence of the first recognition result and generates a corresponding confidence score value;

[0054] If the confidence score value is greater than or equal to the confidence threshold, use the first recognition result as the content of the speech instruction;

[0055] If the confidence score value is less than the confidence threshold, upload the speech instruction to the cloud, so that the cloud recognition model pre-deployed on the cloud generates a corresponding second recognition result according to the speech instruction;

[0056] Obtain the second recognition result and use the second recognition result as the content of the speech instruction.

[0057] In a possible implementation manner, when the type of the target device is a home device with speech recognition function, the control module controls the target device to execute the operation instruction based on the communication method, including:

[0058] Determine the operation that the target device currently needs to execute according to the operation instruction and the executable operation list of the target device;

[0059] Generate a corresponding AI speech instruction according to the operation that the target device currently needs to execute;

[0060] Broadcast the AI speech instruction to the target device, so that the target device recognizes the AI speech instruction according to the built-in device speech recognition model, and then executes the operation instruction, where the device speech recognition model is trained based on a number of AI speech instructions. Description of the Drawings

[0061] Figure 1 : It is a schematic flow chart of a smart home voice control method under heterogeneous resource conditions provided by an embodiment of the present application.

[0062] Figure 2 : It is a schematic flow chart of speech recognition by a speech recognition model in a smart home voice control method under heterogeneous resource conditions provided by an embodiment of the present application.

[0063] Figure 3: Schematic diagram of the setting and update of the confidence threshold in a smart home voice control method under heterogeneous resource conditions provided by an embodiment of the present application

[0064] Figure 4 : Schematic diagram of the update process of the voice recognition model in a smart home voice control method under heterogeneous resource conditions provided by an embodiment of the present application. Detailed implementation manners

[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0066] It should be noted that the step numbers in the text are only for the convenience of explaining specific embodiments and do not serve as a limitation on the execution order of the steps. In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0067] Embodiment 1:

[0068] As Figure 1 shown, Embodiment 1 provides a smart home voice control method under heterogeneous resource conditions, including steps S1 - S4:

[0069] Step S1, obtain the voice command of the user;

[0070] Step S2, input the voice command into a preset voice recognition model so that the voice recognition model recognizes the content of the voice command, and then determine the target device to be controlled by the user and the corresponding operation command;

[0071] Step S3, according to the type of the target device, determine the communication method with the target device, and the type of the target device includes intelligent home devices, home devices with Bluetooth communication or carrier communication functions, home devices with voice recognition functions, and traditional home devices;

[0072] Step S4, control the target device to execute the operation command based on the communication method.

[0073] An embodiment of the present application provides a smart home voice control method under heterogeneous resource conditions. By using a preset voice recognition model to recognize the user's voice command, determine the target device that the user needs to operate and the operation command that the target device needs to execute, and then determine the communication method with the target device according to the type of the target device, so as to control the target device to execute the operation command. Through the use of different communication control methods, the heterogeneous resources of the whole house are integrated in the embodiment of the present application. Whether it is an intelligent home device, a home device with Bluetooth communication or carrier communication function, a home device with voice recognition function, or a traditional home device, all can achieve voice control through the embodiment of the present application, solving the technical problem that the existing smart home voice control cannot take into account the existing home devices, and improving the compatibility of smart home voice control and the user experience.

[0074] In a preferred embodiment, the smart home voice control method provided by the embodiment of the present application is integrated into an intelligent core. The intelligent core receives the user's voice control command and communicates with other home devices according to the voice control command to achieve voice control of the whole house home devices. The typical heterogeneous devices of the smart home system in the embodiment of the present application include an intelligent core, highly intelligent home devices (such as intelligent floor sweeping robots, etc.), home devices with Bluetooth communication function (such as fans, etc.), home devices with carrier communication function (such as curtains, etc.), home devices with simple voice recognition function (such as kettles, etc.), and traditional home devices (such as ordinary lighting lamps, air conditioners, etc.). Among them, the intelligent core serves as the control core of the entire home system, similar to smart speakers such as Xiaoai Speaker, Tmall Genie, and Xiaodu Speaker on the market. It has strong computing power and storage capacity, and can be equipped with a GPU as needed, and can run a voice inference model on it to convert the heard sound into a corresponding command; it has a high-quality sound system with stable frequency and volume, and can convert text into corresponding voice; it has an infrared emission function and can emit infrared signals of different frequencies to control common traditional household appliances (TV, air conditioner); it has Bluetooth Mesh function and can be used as a Bluetooth Mesh gateway to communicate with home devices with Bluetooth function; it is equipped with a broadband carrier chip and can reuse the power line for carrier communication, without the need to lay additional communication lines and without the need to pay communication fees regularly, and can communicate with home devices with carrier communication capabilities; it integrates intelligent control strategies and can automatically switch different combination modes according to user needs (such as performing specified actions at specified times).

[0075] Highly intelligent home devices have certain intelligent functions, can switch different working modes, and can integrate intelligent working algorithms; they have Bluetooth communication functions, can transmit information and receive instructions from the intelligent core via Bluetooth; they integrate carrier chips, can form a carrier network with the intelligent core, communicate via the carrier network, and receive control instructions; they have a certain voice recognition ability and can correctly parse and convert into text or corresponding instructions when the voice is clear and of high quality. They have Bluetooth communication functions, which provide external communication interfaces, control command interfaces, control feedback interfaces, etc. via Bluetooth, and can deeply interact with the intelligent core via Bluetooth.

[0076] Home devices with Bluetooth communication functions have Bluetooth communication functions, which provide external communication interfaces, control command interfaces, control feedback interfaces, etc. via Bluetooth, and can deeply interact with the intelligent core via Bluetooth.

[0077] Home devices with carrier communication functions have power line carrier communication functions, which provide external communication interfaces, control command interfaces, control feedback interfaces, etc. via power line carrier signals, and can deeply interact with the intelligent core via power line carrier.

[0078] Home devices with simple voice recognition functions are limited by weak computing and storage capabilities. Simple voice recognition models can be deployed, and input sources with stable pitch, loudness, and timbre can be recognized. There may be misrecognition for dialects, accented voices, etc.

[0079] Traditional home devices do not have communication components such as Bluetooth and carrier themselves, cannot communicate with the intelligent core, and do not have voice recognition functions. To achieve intelligent control, the socket that powers them needs to be replaced with an intelligent socket with Bluetooth / or carrier communication functions.

[0080] Stronger voice recognition models can be deployed on the intelligent core. However, to obtain higher recognition accuracy and be able to recognize more accents and dialects, the intelligent core needs to be linked to cloud resources. Cloud devices have clustered hardware and virtualization layers, have stronger computing and storage capabilities, integrate GPUs dedicated to AI training and inference, have the ability to dynamically scale up and down, and are capable of deploying models with hundreds of billions of parameters. Performing voice inference in the cloud has higher accuracy. However, transmitting data from the intelligent core to the cloud for recognition and then back to the intelligent core, and then controlling home devices, the entire process takes a long time and the user experience is poor; performing voice inference on the intelligent core and then directly controlling home devices, the latency of the entire environment is low, but the recognition accuracy is relatively low, and there may be misoperations or non-operations, and the user experience is not good enough either; the advantages of both need to be combined to improve the overall user experience.

[0081] In a possible implementation manner, in step S2, the speech recognition model recognizes the content of the speech instruction, including:

[0082] The speech recognition model generates a first recognition result according to the speech instruction;

[0083] The first recognition result is input into a preset recognition result evaluation model, so that the recognition result evaluation model scores the confidence of the first recognition result and generates a corresponding confidence score value;

[0084] If the confidence score value is greater than or equal to the confidence threshold, the first recognition result is used as the content of the speech instruction;

[0085] If the confidence score value is less than the confidence threshold, the speech instruction is uploaded to the cloud, so that the cloud recognition model pre-deployed in the cloud generates a corresponding second recognition result according to the speech instruction;

[0086] The second recognition result is obtained and used as the content of the speech instruction.

[0087] In the embodiment of the present application, a speech recognition method is provided. By performing a confidence score on the first recognition result of the speech recognition model, the accuracy of the first recognition result is evaluated. When the confidence score value is low, the prior art usually responds with a voice broadcast to inform the user to reissue the speech instruction. However, this method cannot guarantee that the user's next speech instruction can be accurately recognized, resulting in the user repeatedly issuing the same speech instruction. Therefore, the embodiment of the present application improves the accuracy of speech recognition through a cloud-edge collaboration method, and deploys a more complex and more accurate cloud recognition model in the cloud to recognize the user's speech instruction. At the same time, the embodiment of the present application takes into account both the recognition response speed and the recognition accuracy. By comparing the confidence score value with the confidence threshold, it is judged whether to start the cloud recognition model for secondary recognition. When the confidence is high, the recognition result of the speech recognition model is directly output. When the confidence is low, secondary recognition is performed through the cloud recognition model, giving full play to the advantages of the high efficiency of the local speech recognition model and the accuracy of the cloud recognition model, and improving the user experience.

[0088] In a preferred embodiment, the specific process of step S2 is as Figure 2 shown and includes:

[0089] 1. A large language model with a large number of parameters is deployed in the cloud, which has a powerful recognition ability;

[0090] 2. The intelligent core deploys a speech model that can be borne by simplified hardware;

[0091] 3. Deploy a program in the intelligent core that can score the results of each recognition. The score range is 0-100, and the meaning of the score is the confidence level of this recognition.

[0092] 4. Select a confidence score (confidence threshold) score as the basis for determining whether to upload to the cloud.

[0093] 5. When the user issues a voice command, the intelligent core first performs recognition. If the confidence level of the recognition result is greater than or equal to score, it directly notifies the home device to execute the command; if the confidence level of the recognition result is less than score, it uploads this voice to the cloud for recognition. After the cloud finishes recognition and returns to the intelligent core, the intelligent core then notifies the home device to execute the command and waits for a reply.

[0094] 6. After receiving the corresponding command, the home device performs relevant operations and replies to the intelligent core. After receiving the reply, the intelligent core ends the wait, otherwise it resends the command after a period of time.

[0095] As described above, performing voice inference in the cloud has a higher accuracy rate, but the entire process takes a longer time; performing voice inference in the intelligent core and then directly controlling the home device has a lower latency in the entire environment, but the recognition accuracy rate is relatively lower. Therefore, it is necessary to combine the advantages of both to improve the overall user experience. Among them, score, as a key parameter determining the running location of voice inference, needs to be updated periodically according to the actual usage situation.

[0096] Therefore, further, the confidence threshold is initialized based on the historical data of other voice recognition models and updated according to the historical recognition data of the current voice recognition model, including:

[0097] Generate an initial confidence threshold based on the historical data of other voice recognition models and set it in the current voice recognition model.

[0098] After a preset update period, obtain the historical recognition data of the current voice recognition model during the update period.

[0099] Analyze the historical recognition data, and count the number of recognition errors of the current voice recognition model and the number of times the inference time of the cloud recognition model exceeds the preset time.

[0100] Generate corresponding first evaluation index and second evaluation index according to the number of recognition errors of the voice recognition model and the number of times the inference time of the cloud recognition model exceeds the preset time.

[0101] Update the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold, and set the confidence threshold in the current voice recognition model.

[0102] The confidence threshold is updated regularly according to the update period.

[0103] Furthermore, updating the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold includes:

[0104] Calculating a comprehensive evaluation index according to a preset index weight, a first evaluation index, and a second evaluation index;

[0105] Increasing or decreasing the initial confidence threshold according to the comprehensive evaluation index to obtain the confidence threshold.

[0106] The embodiment of the present application provides a method for setting and updating a confidence threshold. When the speech recognition model is first deployed, the current speech recognition model has no historical data as a reference. Therefore, an initial confidence threshold is generated according to the historical data of other speech recognition models and set in the current speech recognition model. Since the usage environments and target users of each speech recognition model are different, the initial confidence threshold needs to be updated according to the actual usage situation after being used for a period of time to improve the recognition efficiency and accuracy of the current speech recognition model. Therefore, after the current speech recognition model runs for a period of time, a first evaluation index and a second evaluation index are generated according to the historical recognition data of the current speech recognition model, and then the initial confidence threshold is updated according to the first evaluation index and the second evaluation index. Among them, the first evaluation index is used to evaluate the recognition accuracy of the model, and the first evaluation index is used to evaluate the recognition efficiency of the model. By increasing or decreasing the initial confidence threshold, the updated confidence threshold is more suitable for the current target user, improving the user experience.

[0107] In a preferred embodiment, as Figure 3 shown, the specific process of setting and updating the confidence threshold is as follows:

[0108] 1. Generate an initial score value using the historical data of other home systems and start the home system; usually, multiple home systems autonomously managed by an intelligent core are connected under a cloud system. For example, if the home system is currently being deployed in the 10th household, then the data of the first 9 households can be used for analysis and an initial score value can be generated.

[0109] 2. Collect user usage situations. If misrecognition during inference by the intelligent core causes the home device to malfunction or not act, resulting in the user needing to reissue an instruction, then record that the edge inference disadvantage factor factorEdge (the first evaluation index) is incremented by 1; if inference in the cloud causes an overly long waiting time, resulting in the user reissuing an instruction during the waiting for a response, then record that the cloud inference disadvantage factor factorCloud (the second evaluation index) is incremented by 1;

[0110] 3. Use the weighted sum of factorEdge and factorCloud as the indicator (comprehensive evaluation indicator) to evaluate the quality of the score. The specific formula is as follows:

[0111] factorWeightSum = factorEdge * weightEdge + factorCloud * weightCloud

[0112] Among them, weightEdge + weightCloud = 1. Initially, the default values are weightEdge = 0.5 and weightCloud = 0.5, and the specific values can be adjusted according to user needs. In the initial stage, adjust the value of the score according to the situation of factorEdge and factorCloud at a relatively high frequency (default 1 day) to find a better score that minimizes the weighted sum factorWeightSum of factorEdge and factorCloud.

[0113] 4. After running stably for a period of time, adjust the value of the score at an appropriate frequency (default 1 week);

[0114] 5. When the voice algorithms of the cloud and the intelligent core are updated again, repeat steps 2 - 4 to optimize the voice inference running position.

[0115] In a possible implementation manner, in step S4, when the type of the target device is a home device with Bluetooth communication or carrier communication function, controlling the target device to execute the operation instruction based on the communication method includes:

[0116] Convert the operation instruction into a corresponding control code according to a preset control code correspondence table. The control code correspondence table is constructed based on the executable operation list of the target device. Each executable operation of the target device corresponds to a digital ID with a preset length, and the digital ID is the control code.

[0117] Generate a corresponding control message according to a preset communication protocol and the control code;

[0118] Send the control message to the target device through Bluetooth communication or carrier communication so that the target device executes the operation instruction.

[0119] The embodiments of the present application provide a communication control method for home appliances with Bluetooth communication or carrier communication functions. By using a preset control code correspondence table, the recognized operation instructions are converted into operation codes recognizable by the target device, and then a control message for communicating with the target device is generated, finally converting the user's voice instructions into corresponding control messages, enabling devices without voice recognition to be controlled by voice instructions, and improving the compatibility of smart home voice control and the user experience.

[0120] In a preferred embodiment, the devices capable of communicating with the smart core include home appliances with Bluetooth communication functions such as smart fans, home appliances with carrier communication functions such as smart curtains, and smart sockets (integrated with relays, carrier communication chips, Bluetooth communication modules, infrared control components, etc. according to requirements) such as Xiaomi air conditioner companions. Taking the control of a traditional air conditioner by a smart socket as an example, the interaction and control function implementation methods between such devices and the smart core are described.

[0121] The frame format used for communication is as follows:

[0122]

[0123] Frame start symbol: Identifies the start of a frame of information, with a value of 68H = 01101000B.

[0124] Address field: The address field consists of 6 bytes, each byte being 2-bit BCD code, and the address length can reach 12 decimal digits. Each meter has a unique communication address, which is independent of the physical layer channel. When the length of the address code used is less than 6 bytes, the high bits are filled with "0". When the communication address is 999999999999H, it is a broadcast address, which is only valid for special commands such as broadcast time calibration and broadcast freezing. When sending a broadcast command, no response from the slave station is required. The address field supports abbreviated addressing, that is, starting from several low bits, the remaining high bits are filled with AAH as a wildcard for meter reading operations, and the address field of the slave station's response frame returns the actual communication address. When the address field is transmitted, the low byte comes first and the high byte comes later.

[0125] Control code: The control code is used to identify different control actions.

[0126] Data field length: The data field length L is the number of bytes in the data field. When reading data, L ≤ 200; when writing data, L ≤ 50; L = 0 indicates no data field.

[0127] Data field: The data field includes data identification, password, operator code, data, frame number, etc., and its structure changes with the function of the control code. When transmitting, the sender performs a 33H addition operation on each byte, and the receiver performs a 33H subtraction operation on each byte.

[0128] Checksum: The checksum is the sum of all bytes modulo 256 from the start of the first frame delimiter to the byte before the checksum, i.e., the binary arithmetic sum of each byte, disregarding the overflow value exceeding 256.

[0129] End delimiter: The end delimiter identifies the end of a frame of information, and its value is 16H = 00010110B.

[0130] The following table defines common control commands:

[0131]

[0132]

[0133] Air conditioner on and off: The air conditioner on and off commands are initiated by the intelligent core and transmitted to the smart socket via Bluetooth or carrier wave. The smart socket controls the internal relay to perform corresponding actions, thereby achieving the purpose of controlling the air conditioner on and off.

[0134] Air conditioner temperature setting and mode setting: The air conditioner temperature setting command is initiated by the intelligent core and transmitted to the smart socket via Bluetooth or carrier wave. The smart socket analyzes the received signal and, after internal protocol conversion, emits a corresponding infrared control signal to control the air conditioner to make a response action.

[0135] For home devices with speech recognition capabilities, speech recognition, as a key technology for human-computer interaction, has always been a research hotspot in the field of technology applications. At present, great achievements have been made in speech recognition technology from theoretical research to product development. However, relevant research and application implementation still face great challenges. Neural network models with high accuracy and good effects often require a large amount of computing resources and are huge in scale. However, the computing power and memory of home devices are limited. Even with model compression and acceleration methods such as network pruning, parameter quantization, and knowledge distillation, it is difficult to support different pronunciation habits, accents, and dialects. In a conventional smart home system, the intelligent core has sufficient computing and storage resources and can achieve high-precision recognition for different pronunciation habits, accents, and dialects. Home devices with simple speech recognition capabilities, such as smart kettles, can recognize input sources with stable pitch, loudness, and timbre and are powerless for complex input sources. Human pronunciation will change due to emotions and physical states, while the pronunciation of the intelligent core is completed by corresponding hardware and can be highly consistent. Therefore, the intelligent core can be used as a translator between humans and smart kettles. After a human voice command is issued, the intelligent core with stronger computing power will recognize it, and after obtaining the result, it will issue a control voice to notify the smart kettle to act.

[0136] In a possible implementation manner, when the type of the target device is a home device with voice recognition function, controlling the target device to execute the operation instruction based on the communication method includes:

[0137] Determine the operation that the target device currently needs to execute according to the operation instruction and the list of executable operations of the target device;

[0138] Generate a corresponding AI voice instruction according to the operation that the target device currently needs to execute;

[0139] Broadcast the AI voice instruction to the target device, so that the target device recognizes the AI voice instruction according to the built-in device voice recognition model, and then executes the operation instruction, where the device voice recognition model is trained based on a number of AI voice instructions.

[0140] The embodiment of the present application provides a communication control method for a home device with voice recognition function. Although these home devices have voice recognition function, their computing power and memory are limited. Even if model compression and acceleration methods such as network pruning, parameter quantization, and knowledge distillation are used, it is difficult to support the voice recognition of different users' pronunciation habits, accents, and dialects. And the control center can usually deploy more complex models, with higher voice recognition accuracy than that of home devices. Therefore, the embodiment of the present application improves the voice recognition accuracy of home devices through voice conversion, uses the voice recognition model of the control center to recognize the user's voice instruction instead of these home devices, and then converts it into a corresponding AI voice instruction according to the recognition result. At the same time, a number of AI voice instructions are used to train the device voice recognition model of the home device in advance, so that the home can accurately recognize the AI voice instruction during actual application, improve the voice recognition accuracy of the home device with voice recognition function, and thus improve the user experience.

[0141] In a preferred embodiment, taking a kettle as an example, the control process for a home device with voice recognition function is specifically as follows:

[0142] 1. Collect all the instructions that the intelligent kettle needs to execute, and draw up the voice content that the kettle needs to interact with the intelligent core;

[0143] 2. Let the intelligent core broadcast relevant voice instructions a large number of times and repeatedly and record them as the training data of the intelligent kettle;

[0144] 3. The intelligent kettle uses the obtained training data for training to ensure a high recognition accuracy;

[0145] 4. Establish the correspondence between specific human voices and the instructions issued by the intelligent core. For example, the voices for controlling an intelligent kettle to boil water may be "boil water", "boil the water", "heat", etc. The intelligent core should accurately recognize them and convert them into standard pronunciations for the intelligent kettle to execute;

[0146] 5. Test the control effect multiple times, record the intermediate results of each test, analyze and optimize until the desired accuracy is achieved.

[0147] The most widely used scenario of the smart home system is the home, and the users are relatively fixed. Their pronunciation habits and accents are also relatively fixed. The initial model trained by the large language model is a model with strong generalization ability, but it has not been fine-tuned for a specific home environment. Therefore, after the smart home system has been used for a period of time, we can update and optimize the model according to the collected data and user feedback.

[0148] Therefore, further, the speech recognition model is updated regularly based on a preset time interval, including:

[0149] Collect the historical recognition data of the speech recognition model within the preset time interval. The historical recognition data includes the user's voice commands, the recognition results of the speech recognition model, and the number of times the user repeats the voice;

[0150] Generate a number of training samples after manually annotating the historical recognition data;

[0151] Upload each of the training samples to the cloud so that the cloud can train and optimize the speech recognition model to be updated in the cloud according to each of the training samples, and generate a corresponding updated speech recognition model. Among them, the model parameters of the speech recognition model to be updated are the same as those of the speech recognition model;

[0152] Obtain the updated speech recognition model sent by the cloud, and replace the speech recognition model with the updated speech recognition model for redeployment.

[0153] The most widely used scenario of the smart home system is the home, and the users are relatively fixed. Their pronunciation habits and accents are also relatively fixed. The initially deployed speech recognition model is a model with strong generalization ability, but it has not been fine-tuned for a specific home environment. Therefore, after the smart home system has been used for a period of time, the speech recognition model can be updated and optimized according to the collected data and user feedback. In the embodiments of the present application, the speech recognition model is updated regularly by uploading the historical recognition data to the cloud, so that the local speech recognition model is gradually optimized, improving the accuracy of smart home voice control and thus enhancing the user experience.

[0154] In a preferred embodiment, such as Figure 4As shown, the update process of the speech recognition model is specifically as follows:

[0155] 1. Deploy the trained language model;

[0156] 2. Collect user usage data. If, after the user makes a sound, it is correctly recognized and responded to, this pair of data is used as a positive sample. If, after the user makes a sound, the model fails to recognize the result, this pair of data is used as a key sample for attention and is selected for background analysis. If, after the user makes a sound, the model misrecognizes and causes the user to repeat the sound before it is correctly recognized, the first speech and the first recognition result are used as negative samples, and the second speech and the second recognition result are used as positive samples. If the user repeats the sound without being correctly recognized, the speech and the recognition result are used as key samples for attention and are selected for background analysis. Background analysis includes: using a model deployed on a more powerful cloud server for recognition; having annotators conduct manual analysis and correction to establish the correspondence between the speech and the correct result. Only the key samples for attention are placed in the background analysis, and other samples are used as model training data.

[0157] 3. Sounds generally have context relationships, and people's speech also has certain personal habits. Some words have strong forward and backward associations. Collect the speech over a period of time as a sample;

[0158] 4. After an interval of a specified period (such as 1 week), send the obtained sample data to the cloud to optimize the model;

[0159] 5. Use the collected sound data for simulation testing. After the model accuracy is stably improved compared to before, redeploy it to the production environment;

[0160] 6. Use model compression and acceleration methods such as network pruning, parameter quantization, and knowledge distillation for the new cloud model, and regenerate the model for deployment on the intelligent core.

[0161] Embodiment 2:

[0162] Embodiment 2 provides a smart home voice control system under heterogeneous resource conditions, including an acquisition module 10, a speech recognition module 20, a communication method determination module 30, and a control module 40;

[0163] Among them, the acquisition module 10 is used to acquire the voice commands of the user;

[0164] The speech recognition module 20 is used to input the voice command into a preset speech recognition model, so that the speech recognition model recognizes the content of the voice command, and then determines the target device to be controlled by the user and the corresponding operation command;

[0165] The communication mode determination module 30 is configured to determine the communication mode with the target device according to the type of the target device, where the type of the target device includes intelligent home devices, home devices with Bluetooth communication or carrier communication functions, home devices with voice recognition functions, and traditional home devices;

[0166] The control module 40 is configured to control the target device to execute the operation instruction based on the communication mode.

[0167] In a possible implementation manner, the content of the voice command recognized by the voice recognition model includes:

[0168] The voice recognition model generates a first recognition result according to the voice command;

[0169] Input the first recognition result into a preset recognition result evaluation model, so that the recognition result evaluation model scores the confidence of the first recognition result and generates a corresponding confidence score value;

[0170] If the confidence score value is greater than or equal to the confidence threshold, the first recognition result is used as the content of the voice command;

[0171] If the confidence score value is less than the confidence threshold, the voice command is uploaded to the cloud, so that the cloud recognition model pre-deployed on the cloud generates a corresponding second recognition result according to the voice command;

[0172] Obtain the second recognition result and use the second recognition result as the content of the voice command.

[0173] Furthermore, the confidence threshold is initialized based on the historical data of other voice recognition models and updated according to the historical recognition data of the current voice recognition model, including:

[0174] Generate an initial confidence threshold according to the historical data of other voice recognition models and set it in the current voice recognition model;

[0175] After a preset update period, obtain the historical recognition data of the current voice recognition model during the update period;

[0176] Analyze the historical recognition data, and count the number of recognition errors of the current voice recognition model and the number of times that the inference time of the cloud recognition model exceeds the preset time;

[0177] Generate corresponding first evaluation index and second evaluation index according to the number of recognition errors of the voice recognition model and the number of times that the inference time of the cloud recognition model exceeds the preset time;

[0178] Update the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold, and set the confidence threshold in the current speech recognition model;

[0179] The confidence threshold is updated regularly according to the update period.

[0180] Further, updating the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold includes:

[0181] Calculate a comprehensive evaluation index according to a preset index weight, the first evaluation index, and the second evaluation index;

[0182] Increase or decrease the initial confidence threshold according to the comprehensive evaluation index to obtain the confidence threshold.

[0183] In a possible implementation manner, when the type of the target device is a home device with Bluetooth communication or carrier communication function, the control module 40 controls the target device to execute the operation instruction based on the communication method, including:

[0184] Convert the operation instruction into a corresponding control code according to a preset control code correspondence table, and the control code correspondence table is constructed and generated according to the executable operation list of the target device. Each executable operation of the target device corresponds to a digital ID with a preset length, and the digital ID is the control code.

[0185] Generate a corresponding control message according to a preset communication protocol and the control code;

[0186] Send the control message to the target device through Bluetooth communication or carrier communication, so that the target device executes the operation instruction.

[0187] In a possible implementation manner, when the type of the target device is a home device with speech recognition function, the control module 40 controls the target device to execute the operation instruction based on the communication method, including:

[0188] Determine the operation that the target device currently needs to execute according to the operation instruction and the executable operation list of the target device;

[0189] Generate a corresponding AI voice instruction according to the operation that the target device currently needs to execute;

[0190] Broadcast the AI voice instruction to the target device, so that the target device recognizes the AI voice instruction according to the built-in device speech recognition model, and then executes the operation instruction, where the device speech recognition model is trained based on a number of AI voice instructions.

[0191] Furthermore, the speech recognition model is updated regularly based on a preset time interval, including:

[0192] Collecting historical recognition data of the speech recognition model within the preset time interval, where the historical recognition data includes the user's voice commands, the recognition results of the speech recognition model, and the number of times the user repeats the voice;

[0193] Generating a number of training samples after manually annotating the historical recognition data;

[0194] Uploading each of the training samples to the cloud so that the cloud trains and optimizes the speech recognition model to be updated in the cloud according to each of the training samples, generating a corresponding updated speech recognition model, where the model parameters of the speech recognition model to be updated are the same as those of the speech recognition model;

[0195] Obtaining the updated speech recognition model sent by the cloud, and replacing the speech recognition model with the updated speech recognition model for redeployment.

[0196] The embodiment of the present application provides a smart home voice control system under heterogeneous resource conditions. The voice commands of the user are recognized by a preset speech recognition model to determine the target device that the user needs to operate and the operation command that the target device needs to execute, and then the communication method with the target device is determined according to the type of the target device, so as to control the target device to execute the operation command. By using different communication control methods, the heterogeneous resources of the whole house are integrated in the embodiment of the present application. Whether it is an intelligent home device, a home device with Bluetooth communication or carrier communication function, a home device with speech recognition function, or a traditional home device, voice control can be realized through the embodiment of the present application, solving the technical problem that the existing smart home voice control cannot take into account the existing home devices, and improving the compatibility of smart home voice control and the user experience.

[0197] The more detailed working principle and step flow of this embodiment can, but are not limited to, refer to the relevant records in Embodiment 1.

[0198] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A smart home voice control method under heterogeneous resource conditions, characterized in that: include: Get the user's voice command; Inputting the voice command into a preset voice recognition model so that the voice recognition model recognizes the content of the voice command, thereby determining the target device to be controlled by the user and the corresponding operation command; Determining a communication method with the target device according to the type of the target device; The target device is controlled to execute the operation instruction based on the communication mode.

2. The smart home voice control method under heterogeneous resource conditions as claimed in claim 1, characterized in that: The speech recognition model recognizes the content of the speech instruction, including: The speech recognition model generates a first recognition result according to the speech instruction; Inputting the first recognition result into a preset recognition result evaluation model so that the recognition result evaluation model scores the confidence of the first recognition result and generates a corresponding confidence score value; If the confidence score is greater than or equal to the confidence threshold, taking the first recognition result as the content of the voice instruction; If the confidence score value is less than the confidence threshold, uploading the voice instruction to the cloud, so that the cloud recognition model pre-deployed in the cloud generates a corresponding second recognition result according to the voice instruction; The second recognition result is obtained, and the second recognition result is used as the content of the voice instruction.

3. The smart home voice control method under heterogeneous resource conditions as claimed in claim 2, characterized in that: The confidence threshold is set based on the historical data of other speech recognition models and is updated based on the historical recognition data of the current speech recognition model, including: Generate an initial confidence threshold based on historical data of other speech recognition models and set it on the current speech recognition model; After a preset update cycle, obtaining historical recognition data of the current speech recognition model within the update cycle; Analyze the historical recognition data, and count the number of recognition errors of the current speech recognition model and the number of times the cloud recognition model inference time exceeds a preset time; Generate corresponding first evaluation indicators and second evaluation indicators according to the number of recognition errors of the speech recognition model and the number of times the inference time of the cloud recognition model exceeds a preset time; Updating the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold, and setting the confidence threshold in the current speech recognition model; The confidence threshold is updated periodically according to the update period.

4. The smart home voice control method under heterogeneous resource conditions as claimed in claim 3, characterized in that: Updating the initial confidence threshold according to the first evaluation index and the second evaluation index to obtain the confidence threshold includes: The comprehensive evaluation index is calculated according to the preset index weight, the first evaluation index and the second evaluation index; The initial confidence threshold is increased or decreased according to the comprehensive evaluation index to obtain the confidence threshold.

5. The smart home voice control method under heterogeneous resource conditions as claimed in claim 1, characterized in that: When the type of the target device is a household device with a Bluetooth communication or carrier communication function, the controlling the target device to execute the operation instruction based on the communication mode includes: Convert the operation instruction into a corresponding control code according to a preset control code correspondence table, wherein the control code correspondence table is generated based on a list of executable operations of the target device, and each executable operation of the target device corresponds to a digital ID of a preset length, and the digital ID is the control code; Generate a corresponding control message according to a preset communication protocol and the control code; The control message is sent to the target device via Bluetooth communication or carrier communication, so that the target device executes the operation instruction.

6. The smart home voice control method under heterogeneous resource conditions as claimed in claim 1, characterized in that: When the type of the target device is a household device with a voice recognition function, controlling the target device to execute the operation instruction based on the communication mode includes: Determine the operation currently to be performed by the target device according to the operation instruction and the executable operation list of the target device; Generate a corresponding AI voice command according to the operation currently to be performed by the target device; The AI ​​voice command is broadcast to the target device so that the target device recognizes the AI ​​voice command according to a built-in device voice recognition model and then executes the operation instruction, wherein the device voice recognition model is obtained based on training of several AI voice commands.

7. A smart home voice control method under heterogeneous resource conditions as described in any one of claims 1 to 6, characterized in that: The speech recognition model is regularly updated based on a preset time interval, including: Collecting historical recognition data of the speech recognition model within the preset time interval, the historical recognition data including the user's voice command, the recognition result of the speech recognition model, and the number of times the user repeated utterances; Manually annotating the historical recognition data to generate a number of training samples; Uploading each of the training samples to the cloud, so that the cloud trains and tunes the speech recognition model to be updated in the cloud according to each of the training samples to generate a corresponding updated speech recognition model, wherein the speech recognition model to be updated has the same model parameters as the speech recognition model; The updated speech recognition model sent from the cloud is obtained, and the updated speech recognition model is used to replace the speech recognition model and redeploy it.

8. A smart home voice control system under heterogeneous resource conditions, characterized in that: It includes an acquisition module, a speech recognition module, a communication mode determination module and a control module; Wherein, the acquisition module is used to acquire the user's voice command; The voice recognition module is used to input the voice instruction into a preset voice recognition model so that the voice recognition model recognizes the content of the voice instruction, and then determines the target device to be controlled by the user and the corresponding operation instruction; The communication mode determination module is used to determine the communication mode with the target device according to the type of the target device; The control module is used to control the target device to execute the operation instruction based on the communication mode.

9. The smart home voice control system under heterogeneous resource conditions as claimed in claim 8, characterized in that: The speech recognition model recognizes the content of the speech instruction, including: The speech recognition model generates a first recognition result according to the speech instruction; Inputting the first recognition result into a preset recognition result evaluation model so that the recognition result evaluation model scores the confidence of the first recognition result and generates a corresponding confidence score value; If the confidence score is greater than or equal to the confidence threshold, taking the first recognition result as the content of the voice instruction; If the confidence score value is less than the confidence threshold, uploading the voice instruction to the cloud, so that the cloud recognition model pre-deployed in the cloud generates a corresponding second recognition result according to the voice instruction; The second recognition result is obtained, and the second recognition result is used as the content of the voice instruction.

10. The smart home voice control system under heterogeneous resource conditions as claimed in claim 8, characterized in that: When the type of the target device is a household device with a voice recognition function, the control module controls the target device to execute the operation instruction based on the communication mode, including: Determine the operation currently to be performed by the target device according to the operation instruction and the executable operation list of the target device; Generate a corresponding AI voice command according to the operation currently to be performed by the target device; The AI ​​voice command is broadcast to the target device so that the target device recognizes the AI ​​voice command according to a built-in device voice recognition model and then executes the operation instruction, wherein the device voice recognition model is obtained based on training of several AI voice commands.