Smart home control method and system based on end-side large model

By deploying large end-side models and collaborative environmental monitoring equipment on terminal devices, the problem of traditional smart home control relying on cloud servers is solved, personalized and humanized localized control is achieved, and the convenience and safety of home life are improved.

CN120686642APending Publication Date: 2025-09-23HARBIN SAISI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510840332.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional smart home control relies on cloud servers, which makes it inefficient and vulnerable to attacks, and makes it impossible to achieve personalized and humane localized control.

Method used

A lightweight model is deployed on the terminal device using a large end-side model. Combined with collaborative environmental monitoring equipment, the NPU acceleration component extracts the Mel-frequency cepstral coefficients and zero-crossing rate of voice information, identifies key intent words, and generates executable control instructions to achieve intelligent linkage between devices and abnormal environment detection.

Benefits of technology

Without being in the cloud, privacy protection, personalized and humanized home control are achieved, the convenience and safety of home life are improved, and the coordination and automation level between devices are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120686642A_ABST
    Figure CN120686642A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart home control, and discloses a smart home control method and system based on an end-side large model, and the method comprises the steps: deploying a lightweight end-side large model in a terminal device, collecting the voice information of a target user corresponding to a home control scene, extracting a Mel-frequency cepstral coefficient and a zero-crossing rate of the voice information by using a lightweight end-side large model so as to extract key intention words of the voice information; generating a first executable control instruction of the home control scene; identifying linkage intention features of the key intention words, and generating a second executable control instruction of the home control scene; the method comprises the following steps: acquiring a first executable control instruction and a second executable control instruction, analyzing a user behavior of a target user and an abnormal environment of a home control scene, and generating a target execution control instruction of the home control scene under the condition of the first executable control instruction and the second executable control instruction. And privacy protection, localization, individuation and humanization household intelligent control are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a smart home control method and system based on a terminal-side large model, belonging to the technical field of smart home control. Background Art

[0002] Smart home control refers to the process of centrally managing, remotely controlling, automating, and intelligently decision-making various household electrical appliances, lighting systems, environmental monitoring systems, and security systems using advanced computer, network, and intelligent control technologies. The core goal of smart home control is to improve living comfort, convenience, safety, and energy efficiency while reducing the burden of household chores.

[0003] Traditional smart home control methods typically rely on centralized control systems, using physical switches, remote controls, or wall panels to control various devices in the home. Users must manually operate each device, for example, using separate remote controls for the TV, air conditioner, and lights, or turning lights on and off via wall switches. Devices cannot automatically adjust to environmental changes or user habits, and control signals must be relayed and processed through cloud servers. Any server issues or attacks could affect user home control, resulting in inefficient home control. Summary of the Invention

[0004] The present invention provides a smart home control method and system based on a terminal-side large model, the main purpose of which is to achieve privacy protection, localization, personalization and humanized smart home control without the need for a cloud server.

[0005] To achieve the above objectives, the present invention provides a smart home control method based on a large device model, comprising:

[0006] Determine the scene characteristics of the home control scene to determine the terminal device of the home control scene, deploy a pre-trained lightweight end-side large model in the terminal device, and configure the collaborative environment monitoring device of the terminal device;

[0007] Collecting voice information of the target user corresponding to the home control scenario, and using the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information to extract key intent words from the voice information;

[0008] Based on the key intention word, generating a first executable control instruction for the home control scene;

[0009] Identifying a linkage intention feature of the key intention word, and generating a second executable control instruction for the home control scene based on the linkage intention feature;

[0010] The collaborative environment monitoring device is used to collect environmental data of the home control scene to analyze the user behavior of the target user and the abnormal environment of the home control scene. Based on the user behavior and the abnormal environment, under the conditions of the first executable control instruction and the second executable control instruction, a target execution control instruction of the home control scene is generated, and smart home control of the home control scene is executed based on the target execution control instruction.

[0011] Optionally, configuring the collaborative environment monitoring device of the terminal device includes:

[0012] Analyze the environmental home impact factors of the terminal device corresponding to the home control scene;

[0013] Environmental monitoring equipment for determining said environmental home impact factors;

[0014] Creating a three-dimensional voxel grid of the home control scene;

[0015] Analyzing a voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid;

[0016] determining the installation location of the environmental monitoring device according to the voxel matching coefficient, and establishing a collaborative network of the environmental monitoring device;

[0017] Based on the installation location and the collaborative network, a collaborative environment monitoring device of the terminal device is constructed.

[0018] Optionally, analyzing a voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid includes:

[0019] determining voxel weights of the three-dimensional voxel grid;

[0020] analyzing voxel coverage values ​​of the environmental monitoring device;

[0021] According to the voxel weight and the voxel coverage value, the voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid is calculated using the following formula:

[0022]

[0023] Where SMC represents the voxel matching coefficient of environmental monitoring device i, ω k represents the voxel weight of voxel k in the 3D voxel grid, C i,k represents the voxel coverage value of environmental monitoring device i to voxel k, e represents the exponential function, d i,k represents the physical distance from the environment monitoring device i to the center of the voxel k, r represents the effective radius of the environment monitoring device i, and N represents the number of voxels in the three-dimensional voxel grid.

[0024] Optionally, the extracting the Mel-frequency cepstral coefficients and the zero-crossing rate of the voice information by using the NPU acceleration component corresponding to the lightweight large model on the end side includes:

[0025] Preprocessing the voice information based on the NPU acceleration component to obtain preprocessed voice information;

[0026] Calculating a zero-crossing point of the preprocessed voice information;

[0027] Based on the zero-crossing point, calculating a zero-crossing rate of the preprocessed voice information;

[0028] Pre-emphasize the pre-processed voice information to obtain pre-emphasized voice information;

[0029] Windowing the pre-emphasized voice information to obtain windowed voice information;

[0030] Performing a fast Fourier transform on the windowed speech information to obtain a speech spectrum;

[0031] Extracting the logarithmic energy of the speech spectrum;

[0032] Mel-cepstral coefficients of the speech spectrum are extracted according to the logarithmic energy.

[0033] Optionally, extracting Mel-cepstral coefficients of the speech spectrum according to the logarithmic energy includes:

[0034] Analyzing the frequency domain response and the end-side perception gain factor of the speech spectrum;

[0035] Based on the frequency domain response, the terminal-side perception gain factor, and the logarithmic energy, the Mel-cepstral coefficient of the speech spectrum is calculated using the following formula:

[0036]

[0037] Among them, MFCC represents the Mel cepstral coefficient of the speech spectrum, E c represents the logarithmic energy of the cth Mel filter of the speech spectrum, m represents the Mel cepstral coefficient number, c represents the index of the Mel filter, Q represents the number of Mel filters, P c represents the edge sensing gain factor, θ c represents the frequency domain response of the cth Mel filter of the speech spectrum, cos represents the discrete cosine transform function, and log represents the logarithmic operation function.

[0038] Optionally, extracting key intention words from the voice information includes:

[0039] Analyzing the noise type and noise level of the speech information based on the Mel-frequency cepstral coefficients and the zero-crossing rate of the speech information;

[0040] Denoising the voice information according to the noise type and noise level to obtain denoised voice information;

[0041] Extracting keywords from the denoised voice information and parsing the keywords to obtain parsed keywords;

[0042] The key intention words of the denoised voice information are analyzed according to the parsed keywords.

[0043] Optionally, generating a first executable control instruction for the home control scene based on the key intention word includes:

[0044] Mapping the key intent words to the standard intent of the lightweight end-side large model corresponding to the home control scenario;

[0045] Extracting control parameters and control parameter information of the key intention words;

[0046] Determining a control index of the control parameter based on the control parameter information;

[0047] Based on the control index and the standard intention, a first executable control instruction for the home control scenario is generated.

[0048] Optionally, identifying the linkage intention feature of the key intention word includes:

[0049] Establishing a home device map corresponding to the home control scenario of the key intent word;

[0050] Based on the control parameters of the key intent word, using the home device map to analyze the linkage devices of the key intent word;

[0051] Analyzing the linkage direction of the linkage device to the key intention word;

[0052] According to the linkage direction, the linkage intention feature of the key intention word is determined.

[0053] Optionally, analyzing the user behavior of the target user and the abnormal environment of the home control scenario includes:

[0054] Dividing the environmental data corresponding to the target user into user behavior data and environmental parameter data;

[0055] Analyzing the target user's behavior trajectory and behavior intention based on the user behavior data;

[0056] Determining the user behavior of the target user based on the behavior trajectory and behavior intention;

[0057] The abnormal environment of the home control scene is determined according to the environmental parameter data and a preset environmental parameter threshold.

[0058] In order to solve the above problems, the present invention further provides a smart home control system based on a large device model, the system comprising:

[0059] The device-side large model configuration module is used to determine the scene characteristics of the home control scene to determine the terminal device of the home control scene, deploy the pre-trained lightweight device-side large model in the terminal device, and configure the collaborative environment monitoring device of the terminal device;

[0060] A key intent word recognition module is used to collect voice information of the target user corresponding to the home control scenario, and use the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information to extract the key intent words from the voice information;

[0061] A first control instruction construction module, configured to generate a first executable control instruction for the home control scenario based on the key intention word;

[0062] A second control instruction construction module is used to identify the linkage intention feature of the key intention word and generate a second executable control instruction for the home control scene based on the linkage intention feature;

[0063] A target control instruction construction module is used to use the collaborative environment monitoring device to collect environmental data of the home control scene to analyze the user behavior of the target user and the abnormal environment of the home control scene. Based on the user behavior and the abnormal environment, under the conditions of the first executable control instruction and the second executable control instruction, a target execution control instruction of the home control scene is generated, and smart home control of the home control scene is executed based on the target execution control instruction.

[0064] In the smart home control scenario, by determining the scene features and deploying a lightweight end-side large model, the system can accurately identify the terminal device and configure the collaborative environment monitoring device to achieve efficient environmental data collection. By using the NPU acceleration component to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information, the system can quickly and accurately extract the key intention words, generate the first executable control instruction, and achieve immediate response to the user instruction. Further, by identifying the linkage intention features of the key intention words, the system can generate the second executable control instruction to achieve intelligent linkage between devices and improve the coordination and automation level of home control. The environmental data collected by the collaborative environment monitoring device is not only used to analyze user behavior, but also to detect abnormal environments in home control scenarios to ensure user safety and comfort. Based on the analysis of user behavior and abnormal environments, the system can generate target execution control instructions under the conditions of the first and second executable control instructions to achieve precise control of the smart home. This intelligent control strategy not only improves the convenience and comfort of home life, but also enhances the safety of the home environment through timely detection and response to abnormal environments. Therefore, the present invention can achieve personalized and humanized services for home scene control without being in the cloud. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A schematic diagram of a process flow of a smart home control method based on a device-side large model provided by one embodiment of the present invention;

[0066] Figure 2 A schematic diagram of standard intent mapping for implementing the device-side large model-based smart home control method provided in one embodiment of the present invention;

[0067] Figure 3 A schematic diagram of modules for implementing the device-side large model-based smart home control method provided in one embodiment of the present invention.

[0068] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0069] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0070] The embodiment of the present application provides a smart home control method based on a large end-model. The execution subject of the smart home control method based on the large end-model includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present application. In other words, the smart home control method based on the large end-model can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0071] Example 1:

[0072] Reference Figure 1 FIG2 is a flow chart of a smart home control method based on a large device model according to an embodiment of the present invention. In this embodiment, the smart home control method based on a large device model includes:

[0073] S1. Determine the scene features of the home control scene to determine the terminal device of the home control scene, deploy a pre-trained lightweight end-side large model in the terminal device, and configure the collaborative environment monitoring device of the terminal device.

[0074] It should be explained that home control scenarios refer to specific situations in the home environment that require the smart home system to respond and control based on specific user needs, behaviors, and environmental conditions. Examples include waking up, leaving home, returning home, sleeping, meeting guests, watching a movie, cooking, etc. Scenario features refer to the key attributes and parameters used to describe and distinguish these home control scenarios, such as user behavior characteristics, device status characteristics, and environmental parameter characteristics.

[0075] It should be explained that the terminal device refers to an intelligent hardware device that can directly interact with the user or directly control various devices in the home environment (such as lights, air conditioners, curtains, TVs, etc.). The present invention adopts Espressif's ESP32-P4 series MCU equipped with a RISC-V dual-core processor. The present invention is based on RISC-V, Linux, and ARM hardware platforms, and realizes whole-house voice interaction, environmental adjustment, security monitoring, multi-device collaboration and other functions through a large end-side model. All data processing and decision-making are completed locally without cloud dependence.

[0076] The present invention deploys a pre-trained lightweight end-side large model in the terminal device to quickly process user voice commands, improve the response efficiency of home control, and enable the smart home system to continue operating even when there is no external network or cloud server failure or cloud server attack, thereby achieving privacy protection, localized services, low cost, high efficiency, and stability. The lightweight end-side large model refers to an artificial intelligence model deployed locally on the terminal device, and its parameter scale is generally smaller than the cloud-side large model (for example, 7 billion parameters, while the cloud-side model, such as GPT-4, has 1.8 trillion parameters).

[0077] The collaborative environment monitoring device configured with the terminal device of the present invention can collect environmental data of home control scenes as a basis for optimizing user instructions, thereby improving the user experience.

[0078] In detail, the collaborative environment monitoring device configured with the terminal device includes:

[0079] Analyze the environmental home impact factors of the terminal device corresponding to the home control scene;

[0080] Environmental monitoring equipment for determining said environmental home impact factors;

[0081] Creating a three-dimensional voxel grid of the home control scene;

[0082] Analyzing a voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid;

[0083] determining the installation location of the environmental monitoring device according to the voxel matching coefficient, and establishing a collaborative network of the environmental monitoring device;

[0084] Based on the installation location and the collaborative network, a collaborative environment monitoring device of the terminal device is constructed.

[0085] Among them, the environmental home influencing factors refer to various physical quantities that can affect the state of the home environment, such as temperature, humidity, light intensity, etc.; the environmental monitoring equipment refers to a device used to perceive and measure the above-mentioned environmental home influencing factors; the three-dimensional voxel grid refers to a grid that divides the home environment into multiple small cubes (voxels); the voxel matching coefficient refers to the degree of matching between the environmental monitoring device and each voxel in the three-dimensional voxel grid; the installation position refers to the installation position of the environmental monitoring device in the three-dimensional voxel grid; the collaborative network refers to a network composed of multiple environmental monitoring devices, which are interconnected and work in collaboration through wireless communication (such as Zigbee, Wi-Fi, Bluetooth, etc.); the collaborative environmental monitoring device refers to a set of devices that can work in collaboration with other devices.

[0086] Optionally, the three-dimensional voxel grid for establishing the home control scene may be generated by scanning the home control scene using SLAM technology to generate a voxel grid with a side length of 20 cm.

[0087] Furthermore, the analyzing the voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid includes:

[0088] determining voxel weights of the three-dimensional voxel grid;

[0089] analyzing voxel coverage values ​​of the environmental monitoring device;

[0090] According to the voxel weight and the voxel coverage value, the voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid is calculated using the following formula:

[0091]

[0092] Where SMC represents the voxel matching coefficient of environmental monitoring device i, ω k represents the voxel weight of voxel k in the 3D voxel grid, C i,k represents the voxel coverage value of environmental monitoring device i to voxel k, e represents the exponential function, d i,k represents the physical distance from the environment monitoring device i to the center of the voxel k, r represents the effective radius of the environment monitoring device i, and N represents the number of voxels in the three-dimensional voxel grid.

[0093] Among them, the voxel weight refers to the importance or influence of each voxel in the three-dimensional voxel grid, which can be dynamically set in combination with the scene semantics (such as the smoke monitoring weight of the kitchen > the living room). The voxel coverage value refers to the coverage degree of the environmental monitoring equipment on a specific voxel. The physical distance refers to the straight-line distance from the environmental monitoring equipment to the center of the voxel. The effective radius refers to the maximum distance that the environmental monitoring equipment can effectively monitor.

[0094] S2. Collect the voice information of the target user corresponding to the home control scenario, and use the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information to extract the key intent words of the voice information.

[0095] It should be explained that the target user refers to a person who uses the smart home system, and the voice information refers to the voice data of the target user collected through a voice input device (such as a microphone).

[0096] The present invention utilizes the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information, thereby improving the processing efficiency of the voice information.

[0097] In detail, the extracting of the Mel-frequency cepstral coefficients and zero-crossing rate of the speech information by using the NPU acceleration component corresponding to the lightweight large model on the end side includes:

[0098] Preprocessing the voice information based on the NPU acceleration component to obtain preprocessed voice information;

[0099] Calculating a zero-crossing point of the preprocessed voice information;

[0100] Based on the zero-crossing point, calculating a zero-crossing rate of the preprocessed voice information;

[0101] Pre-emphasize the pre-processed voice information to obtain pre-emphasized voice information;

[0102] Windowing the pre-emphasized voice information to obtain windowed voice information;

[0103] Performing a fast Fourier transform on the windowed speech information to obtain a speech spectrum;

[0104] Extracting the logarithmic energy of the speech spectrum;

[0105] Mel-cepstral coefficients of the speech spectrum are extracted according to the logarithmic energy.

[0106] Among them, the zero crossing point refers to the instantaneous point at which the speech signal changes from positive to negative or from negative to positive in the time domain, the zero crossing rate refers to the number of times the speech signal crosses zero in unit time, the pre-emphasized speech information refers to the speech information after high-pass filtering is performed on the speech signal to enhance the amplitude of the high-frequency part, the windowed speech information refers to the speech information after the pre-emphasized speech signal is divided into short time frames and a window function is applied to each frame, the speech spectrum refers to the frequency domain table obtained by performing fast Fourier transform (FFT) on the windowed speech information, the logarithmic energy refers to the logarithm of the amplitude of the speech spectrum, and the Mel cepstral coefficient refers to the characteristic coefficient obtained by passing the logarithmic energy of the speech spectrum through a Mel filter bank and then performing discrete cosine transform (DCT).

[0107] Optionally, windowing the pre-emphasized speech information to obtain the windowed speech information may be achieved by using a Hamming window.

[0108] Furthermore, extracting the Mel-cepstral coefficients of the speech spectrum according to the logarithmic energy includes:

[0109] Analyzing the frequency domain response and the end-side perception gain factor of the speech spectrum;

[0110] Based on the frequency domain response, the terminal-side perception gain factor, and the logarithmic energy, the Mel-cepstral coefficient of the speech spectrum is calculated using the following formula:

[0111]

[0112] Among them, MFCC represents the Mel cepstral coefficient of the speech spectrum, E c represents the logarithmic energy of the cth Mel filter of the speech spectrum, m represents the Mel cepstral coefficient number, c represents the index of the Mel filter, Q represents the number of Mel filters, P c represents the edge sensing gain factor, θ c represents the frequency domain response of the cth Mel filter of the speech spectrum, cos represents the discrete cosine transform function, and log represents the logarithmic operation function.

[0113] The frequency domain response refers to the response characteristics of the Mel filter in the frequency domain. The end-side perception gain factor refers to a parameter that takes into account the perception characteristics of the end-side device when calculating the Mel cepstral coefficients and is dynamically adjusted by the microphone signal-to-noise ratio (SNR). The Mel cepstral coefficient sequence number refers to the index of the Mel cepstral coefficient, indicating the number of the Mel cepstral coefficient. Typically, there are 13 Mel cepstral coefficients. The discrete cosine transform function refers to a transform algorithm that transforms a signal from the time domain to the frequency domain. The logarithmic operation function refers to an algorithm that performs a logarithmic transform on the output energy of the Mel filter.

[0114] The present invention extracts key intention words from the voice information and can accurately identify the user's intention, thereby improving the accuracy of home control.

[0115] Specifically, extracting the key intention words of the voice information includes:

[0116] Analyzing the noise type and noise level of the speech information based on the Mel-frequency cepstral coefficients and the zero-crossing rate of the speech information;

[0117] Denoising the voice information according to the noise type and noise level to obtain denoised voice information;

[0118] Extracting keywords from the denoised voice information and parsing the keywords to obtain parsed keywords;

[0119] The key intention words of the denoised voice information are analyzed according to the parsed keywords.

[0120] The noise type refers to the different types of background noise mixed into the voice signal. Common noise types include environmental noise (such as traffic noise and crowd noise), equipment noise (such as fan noise and electromagnetic interference), and white noise. The noise level refers to the intensity or degree of interference of the noise. The denoised voice information refers to the voice signal that has been denoised. The parsed keywords refer to the meaningful key words extracted from the denoised voice information. The key intent words refer to the core words that can accurately express the user's intent.

[0121] Optionally, extracting keywords from the denoised voice information and parsing the keywords to obtain parsed keywords can be achieved through natural language processing (NLP) and semantic role labeling.

[0122] S3. Based on the key intention word, generate a first executable control instruction for the home control scenario.

[0123] The present invention generates the first executable control instruction for the home control scenario based on the key intention words, which can convert the key intention understood from the natural language voice instruction into a specific and structured control instruction that can be understood and executed by the smart home platform.

[0124] In detail, the generating of the first executable control instruction of the home control scene based on the key intention word includes:

[0125] Mapping the key intent words to the standard intent of the lightweight end-side large model corresponding to the home control scenario;

[0126] Extracting control parameters and control parameter information of the key intention words;

[0127] Determining a control index of the control parameter based on the control parameter information;

[0128] Based on the control index and the standard intention, a first executable control instruction for the home control scenario is generated.

[0129] Among them, the standard intent refers to a set of standardized and normalized operation instruction types pre-defined within the smart home control system. They represent the core actions that the user wants to perform, such as "turn on the light", "turn on the light", "light on", "turn off the air conditioner", "air conditioner off" and other actions. The control parameters refer to the parameters that need to be controlled under the standard intent, such as position, increase, decrease, open and other parameters. The control parameter information refers to the relevant information describing the control parameters, such as increasing the temperature by a little, opening the curtains halfway, and the control index refers to the quantified index of the control parameter, such as increasing the air conditioner temperature in the living room by 5 degrees. The first executable control instruction refers to the final command generated according to the determined standard intent and the corresponding control index, which can be directly sent to the smart home device or control gateway for executing the operation.

[0130] Optionally, the control parameter information-based control indicator can be determined by deriving a corresponding control indicator based on the identified control parameter information using a predefined rule or decision tree. For example, if "temperature" and "26 degrees" are identified, the control indicator "target_t26" is derived based on the rule.

[0131] Reference Figure 2 As shown, a schematic diagram of standard intent mapping for implementing the smart home control method based on the end-side large model provided by one embodiment of the present invention is provided: wherein, the semantic vector encoding refers to the process of converting the input key intent words into a mathematical representation, and the matching standard intent library refers to the process of comparing and matching the input after semantic vector encoding with predefined standard intents, wherein the standard intent library refers to a predefined set of standard intents, and each intent has its corresponding semantic vector representation.

[0132] S4. Identify the linkage intention feature of the key intention word, and generate a second executable control instruction for the home control scene based on the linkage intention feature.

[0133] The present invention can effectively identify the linkage intention features of the key intention words by identifying the linkage intention features of the key intention words, and realize intelligent linkage control of smart home devices.

[0134] Specifically, identifying the linkage intention features of the key intention words includes:

[0135] Establishing a home device map corresponding to the home control scenario of the key intent word;

[0136] Based on the control parameters of the key intent word, using the home device map to analyze the linkage devices of the key intent word;

[0137] Analyzing the linkage direction of the linkage device to the key intention word;

[0138] According to the linkage direction, the linkage intention feature of the key intention word is determined.

[0139] Among them, the home device map refers to a network model of the relationship and interaction between home devices, the linked devices refer to other home devices related to the key intent words that need to work together to realize the user's intention, the linkage direction refers to the direction of control or interaction between devices, and the linkage intention feature refers to a feature derived based on the analysis of key intent words, linked devices and linkage directions, which describes how the user's intention is realized through device linkage. For example, for the key intent word "turn on theater mode", the linkage intention features may include "turn on the TV power", "turn up the audio volume" and "dim the lights".

[0140] Optionally, the establishment of the home device graph corresponding to the home control scenario of the key intent word can be constructed by using graph database technology to use nodes, edges and attributes to represent and store home device data.

[0141] The present invention generates the second executable control instruction of the home control scene based on the linkage intention feature to realize the linkage control of the home control scene and improve the user experience. The second executable control instruction refers to a control instruction that optimizes the linkage of the first executable control instruction according to the intention analysis. For example, the instruction contains "living room" + "light", and automatically associates the same area device "curtain" as a linkage candidate. Exemplarily, the effective linkage of the first executable control instruction and the second executable control instruction can be achieved through multi-device collaborative control (Linux hub + ARM-M4), wherein the first executable control instruction can be lighting control: the Linux hub sends a PWM dimming instruction (brightness 50%) to the ARM-M4 lamp through WIFIMesh. The second executable control instruction can be air conditioning control: the air conditioning mode (cooling + automatic wind speed) is adjusted through the Matter protocol, and the response delay is <100ms.

[0142] S5. Use the collaborative environment monitoring device to collect environmental data of the home control scene to analyze the user behavior of the target user and the abnormal environment of the home control scene. Based on the user behavior and the abnormal environment, under the conditions of the first executable control instruction and the second executable control instruction, generate a target execution control instruction for the home control scene, and execute smart home control of the home control scene based on the target execution control instruction.

[0143] It should be explained that the environmental data refers to the environmental parameters of the home control scene, such as temperature, humidity, gas, user behavior and other data.

[0144] The present invention analyzes the user behavior and abnormal environment of the target user and can serve as a basis for optimizing control instructions.

[0145] In detail, the analyzing the user behavior of the target user and the abnormal environment of the home control scene includes:

[0146] Dividing the environmental data corresponding to the target user into user behavior data and environmental parameter data;

[0147] Analyzing the target user's behavior trajectory and behavior intention based on the user behavior data;

[0148] Determining the user behavior of the target user based on the behavior trajectory and behavior intention;

[0149] The abnormal environment of the home control scene is determined according to the environmental parameter data and a preset environmental parameter threshold.

[0150] The user behavior data refers to data related to the target user's interactions with devices and activity patterns in the smart home environment. The environmental parameter data refers to real-time data about the state of the home environment collected by sensors and other devices. This may include temperature, humidity, light intensity, air quality, noise level, etc. The behavior trajectory refers to the user's movement path and activity record in the home environment within a specific time period. The behavioral intention refers to the purpose or goal of the user performing a certain action or series of actions. The user behavior refers to the specific operations and activity patterns exhibited by the user in the smart home environment. The abnormal environment refers to certain parameters of the home environment exceeding the preset safety or comfort thresholds. For example, the temperature is too high or too low, the air quality is poor, the noise is too loud, etc.

[0151] For example, in a home control scenario, air conditioners can dynamically adjust the temperature based on user habits and indoor and outdoor temperature and humidity. This can be achieved through differential privacy and edge collaborative reasoning: the device locally learns user preferences (e.g., a preference for 25°C at night) and only uploads model parameters (not raw data) to the cloud for aggregation. Multi-sensor fusion: The device-side model integrates temperature and humidity sensors and human infrared data to infer optimal settings in real time.

[0152] Alternatively, user behavior data can be detected in real time at 30fps using a lightweight large-scale model on the edge (e.g., Rockchip RK3588) using YOLO-Nano or MobileViT. This allows for real-time identification of strangers and unusual behavior (e.g., falls and intrusions).

[0153] The present invention is based on the user behavior and the abnormal environment. Under the conditions of the first executable control instruction and the second executable control instruction, the target execution control instruction of the home control scene is generated to realize personalized home control under the conditions of satisfying the user's specific instructions. Among them, the target execution control refers to a control instruction that further improves the user experience under the first executable control instruction and the second executable control instruction. For example, when the user issues an instruction to turn on the TV, it is detected that the user is moving towards the toilet, then the target execution control instruction is "turn on the TV", "close the curtains", "turn on the toilet light", "lower the living room light" and other instructions. In the smart home control scenario, by determining the scene characteristics and deploying a lightweight end-side large model, the system can accurately identify the terminal device and configure collaborative environmental monitoring equipment to achieve efficient environmental data collection. By using the NPU acceleration component to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information, the system can quickly and accurately extract the key intent words, generate the first executable control instruction, and realize Instant response to user instructions is now possible. Furthermore, by identifying the linkage intention features of key intention words, the system can generate a second executable control instruction to achieve intelligent linkage between devices, improve the coordination and automation level of home control, and use the environmental data collected by the collaborative environmental monitoring equipment not only to analyze user behavior, but also to detect abnormal environments in home control scenes to ensure user safety and comfort. Based on the analysis of user behavior and abnormal environments, the system can generate target execution control instructions under the conditions of the first and second executable control instructions to achieve precise control of smart homes. This intelligent control strategy not only improves the convenience and comfort of home life, but also enhances the safety of the home environment through timely detection and response to abnormal environments. Therefore, the present invention can achieve personalized and humanized services for home scene control without being in the cloud.

[0154] Example 2:

[0155] like Figure 3 The figure shows a functional module diagram of a smart home control system based on a large end-side model according to the present invention.

[0156] The device-side large model-based smart home control system 300 described in the present invention can be installed in an electronic device. Depending on the functionality implemented, the device-side large model-based smart home control system may include a device-side large model configuration module 301, a key intent word recognition module 302, a first control instruction construction module 303, a second control instruction construction module 304, and a target control instruction construction module 305. The module described in the present invention, also referred to as a unit, refers to a series of computer program segments that can be executed by an electronic device processor and can perform fixed functions, and is stored in the memory of the electronic device.

[0157] In the embodiment of the present invention, the functions of each module / unit are as follows:

[0158] The device-side large model configuration module 301 is used to determine the scene characteristics of the home control scene to determine the terminal device of the home control scene, deploy the pre-trained lightweight device-side large model in the terminal device, and configure the collaborative environment monitoring device of the terminal device;

[0159] The key intent word recognition module 302 is used to collect voice information of the target user corresponding to the home control scenario, and use the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information to extract the key intent words from the voice information;

[0160] The first control instruction construction module 303 is used to generate a first executable control instruction for the home control scenario based on the key intention word;

[0161] The second control instruction construction module 304 is used to identify the linkage intention feature of the key intention word and generate a second executable control instruction for the home control scene based on the linkage intention feature;

[0162] The target control instruction construction module 305 is used to use the collaborative environment monitoring device to collect environmental data of the home control scene to analyze the user behavior of the target user and the abnormal environment of the home control scene. Based on the user behavior and the abnormal environment, under the conditions of the first executable control instruction and the second executable control instruction, a target execution control instruction of the home control scene is generated, and smart home control of the home control scene is executed based on the target execution control instruction.

[0163] In detail, the modules in the smart home control system 200 based on the end-side large model in the embodiment of the present invention are used in the same manner as above. Figure 1 The same technical means are used as the smart home control method based on the end-side large model described in , and can produce the same technical effects, so I will not go into details here.

[0164] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A smart home control method based on a large device model, characterized in that: The method comprises: Determine the scene characteristics of the home control scene to determine the terminal device of the home control scene, deploy a pre-trained lightweight end-side large model in the terminal device, and configure the collaborative environment monitoring device of the terminal device; Collecting voice information of the target user corresponding to the home control scenario, and using the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information to extract key intent words from the voice information; Based on the key intention word, generating a first executable control instruction for the home control scene; Identifying a linkage intention feature of the key intention word, and generating a second executable control instruction for the home control scene based on the linkage intention feature; The collaborative environment monitoring device is used to collect environmental data of the home control scene to analyze the user behavior of the target user and the abnormal environment of the home control scene. Based on the user behavior and the abnormal environment, under the conditions of the first executable control instruction and the second executable control instruction, a target execution control instruction of the home control scene is generated, and smart home control of the home control scene is executed based on the target execution control instruction.

2. The smart home control method based on the terminal-side large model according to claim 1, characterized in that: The collaborative environment monitoring device configured with the terminal device includes: Analyze the environmental home impact factors of the terminal device corresponding to the home control scene; Environmental monitoring equipment for determining said environmental home impact factors; Creating a three-dimensional voxel grid of the home control scene; Analyzing a voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid; determining the installation location of the environmental monitoring device according to the voxel matching coefficient, and establishing a collaborative network of the environmental monitoring device; Based on the installation location and the collaborative network, a collaborative environment monitoring device of the terminal device is constructed.

3. The smart home control method based on the terminal-side large model according to claim 2, characterized in that: The analyzing the voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid includes: determining voxel weights of the three-dimensional voxel grid; analyzing voxel coverage values ​​of the environmental monitoring device; According to the voxel weight and the voxel coverage value, the voxel matching coefficient of the environmental monitoring device in the three-dimensional voxel grid is calculated using the following formula: Where SMC represents the voxel matching coefficient of environmental monitoring device i, ω k represents the voxel weight of voxel k in the 3D voxel grid, C i,k represents the voxel coverage value of environmental monitoring device i to voxel k, e represents the exponential function, d i,k represents the physical distance from the environment monitoring device i to the center of the voxel k, r represents the effective radius of the environment monitoring device i, and N represents the number of voxels in the three-dimensional voxel grid.

4. The smart home control method based on the terminal-side large model according to claim 3, characterized in that: The extracting of the Mel-frequency cepstral coefficients and the zero-crossing rate of the voice information by using the NPU acceleration component corresponding to the lightweight large end-side model includes: Preprocessing the voice information based on the NPU acceleration component to obtain preprocessed voice information; Calculating a zero-crossing point of the preprocessed voice information; Based on the zero-crossing point, calculating a zero-crossing rate of the preprocessed voice information; Pre-emphasize the pre-processed voice information to obtain pre-emphasized voice information; Windowing the pre-emphasized voice information to obtain windowed voice information; Performing a fast Fourier transform on the windowed speech information to obtain a speech spectrum; Extracting the logarithmic energy of the speech spectrum; Mel-cepstral coefficients of the speech spectrum are extracted according to the logarithmic energy.

5. The smart home control method based on the terminal-side large model according to claim 4 is characterized in that: The step of extracting Mel-cepstral coefficients of the speech spectrum according to the logarithmic energy includes: Analyzing the frequency domain response and the end-side perception gain factor of the speech spectrum; Based on the frequency domain response, the terminal-side perception gain factor, and the logarithmic energy, the Mel-cepstral coefficient of the speech spectrum is calculated using the following formula: Among them, MFCC represents the Mel cepstral coefficient of the speech spectrum, E c represents the logarithmic energy of the cth Mel filter of the speech spectrum, m represents the Mel cepstral coefficient number, c represents the index of the Mel filter, Q represents the number of Mel filters, P c represents the edge sensing gain factor, θ c represents the frequency domain response of the cth Mel filter of the speech spectrum, cos represents the discrete cosine transform function, and log represents the logarithmic operation function.

6. The smart home control method based on the terminal-side large model according to claim 5, characterized in that: The extracting key intention words from the voice information includes: Analyzing the noise type and noise level of the speech information based on the Mel-frequency cepstral coefficients and the zero-crossing rate of the speech information; Denoising the voice information according to the noise type and noise level to obtain denoised voice information; Extracting keywords from the denoised voice information and parsing the keywords to obtain parsed keywords; The key intention words of the denoised voice information are analyzed according to the parsed keywords.

7. The smart home control method based on the terminal-side large model according to claim 6, characterized in that: The generating, based on the key intention word, a first executable control instruction for the home control scene includes: Mapping the key intent words to the standard intent of the lightweight end-side large model corresponding to the home control scenario; Extracting control parameters and control parameter information of the key intention words; Determining a control index of the control parameter based on the control parameter information; Based on the control index and the standard intention, a first executable control instruction for the home control scenario is generated.

8. The smart home control method based on the terminal-side large model according to claim 7, characterized in that: The identification of the linkage intention feature of the key intention word includes: Establishing a home device map corresponding to the home control scenario of the key intent word; Based on the control parameters of the key intent word, using the home device map to analyze the linkage devices of the key intent word; Analyzing the linkage direction of the linkage device to the key intention word; According to the linkage direction, the linkage intention feature of the key intention word is determined.

9. The smart home control method based on the terminal-side large model according to claim 8, characterized in that: The analyzing the user behavior of the target user and the abnormal environment of the home control scene includes: Dividing the environmental data corresponding to the target user into user behavior data and environmental parameter data; Analyzing the target user's behavior trajectory and behavior intention based on the user behavior data; Determining the user behavior of the target user based on the behavior trajectory and behavior intention; The abnormal environment of the home control scene is determined according to the environmental parameter data and a preset environmental parameter threshold.

10. A smart home control system based on a large end-to-end model, characterized in that: The system comprises: The device-side large model configuration module is used to determine the scene characteristics of the home control scene to determine the terminal device of the home control scene, deploy the pre-trained lightweight device-side large model in the terminal device, and configure the collaborative environment monitoring device of the terminal device; A key intent word recognition module is used to collect voice information of the target user corresponding to the home control scenario, and use the NPU acceleration component corresponding to the lightweight end-side large model to extract the Mel-frequency cepstral coefficients and zero-crossing rate of the voice information to extract the key intent words from the voice information; A first control instruction construction module, configured to generate a first executable control instruction for the home control scenario based on the key intention word; A second control instruction construction module is used to identify the linkage intention feature of the key intention word and generate a second executable control instruction for the home control scene based on the linkage intention feature; A target control instruction construction module is used to use the collaborative environment monitoring device to collect environmental data of the home control scene to analyze the user behavior of the target user and the abnormal environment of the home control scene. Based on the user behavior and the abnormal environment, under the conditions of the first executable control instruction and the second executable control instruction, a target execution control instruction of the home control scene is generated, and smart home control of the home control scene is executed based on the target execution control instruction.

Citation Information

Cited By

  • Smart home control method and system based on voice recognition

    CN121768389A