Smart home intelligent interaction system and method based on big data

By introducing technical means of multimodal interaction, regional command differential adaptation and intelligent decision-making in the smart home system, the problems of insufficient single mode, regional adaptation and intelligent decision-making capabilities of existing smart home interaction technologies are solved, and a more efficient, convenient and intelligent user interaction experience is achieved.

CN120223455AInactive Publication Date: 2025-06-27CHONGQING UNIV OF EDUCATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510369158.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing smart home interaction technology has single modal interaction limitations, the inability to adapt to regional command differences, and the lack of intelligent decision-making capabilities that comprehensively consider time, environment and multi-user scenarios.

Method used

Using a smart home intelligent interaction system based on big data, multi-modal interaction, regional command differential adaptation and intelligent decision-making are achieved through the collaborative work of cloud servers, user terminals, routing nodes and smart home nodes. Specific measures include: multimodal data acquisition, load balancing allocation, regional instruction learning database construction, and the application of comprehensive instruction recognition models.

Benefits of technology

It realizes the convenience of multimodal interaction, precise adaptation of regional command differences, and real-time and accuracy of intelligent decision-making, improving user life experience and the efficiency of smart home use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223455A_ABST
    Figure CN120223455A_ABST
Patent Text Reader

Abstract

The invention relates to the field of man-machine interaction, and particularly discloses a smart home intelligent interaction system based on big data, which comprises a cloud server, a user terminal, a routing node and a plurality of smart home nodes, the smart home nodes are distributed in a home environment and used for collecting environmental parameters and original parameters input by a user in real time and sending the environmental parameters and the original parameters to the routing node, and the environmental parameters comprise geographical location information; the routing node is used for determining the routing node or a certain smart home node as an identification main body of the instruction identification according to a preset distribution rule after receiving the original parameters; and the identification main body obtains the environment parameters and the original parameters and obtains an instruction identification result and an instruction matching rate according to a preset instruction identification model. By adopting the technical scheme of the invention, multi-modal interaction can be fused, regional instruction differences can be accurately adapted, and intelligent decision making can be carried out in combination with multi-key information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction, and in particular to a smart home intelligent interaction system and method based on big data. Background Art

[0002] With the rapid development of science and technology, the field of smart home has made great progress, and many interactive systems have emerged to provide users with a convenient and comfortable life experience. However, the current smart home interactive technology still has many limitations.

[0003] In terms of interaction methods, most traditional smart home systems rely on a single mode, such as voice recognition technology. This results in a significant reduction in the convenience of interaction when users are in special situations in actual use scenarios. For example, when users have busy hands and cannot operate voice devices or free their hands for touch operations, a system that only supports voice interaction cannot meet the needs; in another example, in scenarios that require a quiet environment, voice commands will cause trouble, and there is a lack of convenient alternative interaction methods.

[0004] From the perspective of command recognition accuracy, the existing system fails to fully consider the impact of regional differences on user command expression. Due to differences in culture, customs, and language habits in different regions, users use different command words when controlling smart home devices. However, most interactive systems use a unified command vocabulary, which makes the first recognition accuracy of the same command in different regions low, requiring users to try repeatedly, which greatly affects the user experience.

[0005] In addition, existing smart home systems are also lacking in intelligent decision-making capabilities. They often process single commands in isolation, lacking comprehensive consideration of time, environment, and multi-user communication scenarios. In scenarios such as light changes in the early morning and people's activities late at night, they cannot proactively and intelligently respond to user needs; in occasions such as family gatherings where multiple people communicate frequently, they cannot quickly integrate complete command intentions from fragmented keyword exchanges, making it difficult to achieve one-click intelligent linkage scene creation.

[0006] To sum up, the existing smart home interaction technology is in urgent need of improvement. There is an urgent need for a smart home intelligent interaction system and method based on big data that can integrate multimodal interaction, accurately adapt to regional command differences, and combine multiple key information for intelligent decision-making, so as to meet the growing demand for smart home use and improve the quality of life of users. Summary of the invention

[0007] The present invention provides a smart home intelligent interaction system and method based on big data, which can integrate multi-modal interaction, accurately adapt to regional command differences, and make intelligent decisions based on multiple key information.

[0008] To solve the above technical problems, the present application provides the following technical solutions:

[0009] A smart home intelligent interaction system based on big data, comprising: a cloud server, a user terminal, a routing node, and a plurality of smart home nodes;

[0010] The smart home nodes are distributed in the home environment, and are used for real-time collecting environmental parameters and original parameters input by users, and sending the environmental parameters and original parameters to the routing node, where the environmental parameters include geographical location information;

[0011] The routing node is configured to, after receiving the original parameters, determine, according to a preset allocation rule, the routing node or a certain smart home node as the recognition subject for the current instruction recognition; the recognition subject obtains the environmental parameters and original parameters and obtains an instruction recognition result and an instruction matching rate according to a preset instruction recognition model; if the instruction matching rate is lower than a preset value, an instruction confirmation request is actively sent to the user terminal, and at the same time, an assistance recognition request is sent to the cloud server. If, within a preset time, an instruction confirmation message or a modified instruction recognition result input by the user through the user terminal is received, the instruction confirmation message or the modified instruction recognition result is sent to the corresponding smart home node; then, the environmental parameters, original parameters, and instruction recognition result in the current recognition are combined into first object information and sent to the cloud server;

[0012] The cloud server is configured to process the first object information, and construct a regionalized instruction learning database with the geographical location information and the instruction matching rate therein as indexes; and after receiving the assistance recognition request, perform a quick match of the instruction recognition result by indexing with the geographical location information and the instruction matching rate corresponding to the assistance request, and then feedback the successfully matched instruction recognition result to the routing node;

[0013] Wherein, if the routing node does not receive a user feedback within a preset time, the instruction recognition result sent by the cloud server is directly sent to the corresponding smart home node.

[0014] Further, the preset allocation rule is that the routing node monitors its own first load condition in real time at a frequency of at least once per second, and simultaneously sends a load query instruction to each smart home node. After receiving the instruction, each smart home node feeds back its own second load condition. The routing node constructs a load matrix based on the first load condition and the second load condition. When the original parameter arrives, if the load of the routing node meets the first preset condition and there is no idle smart home node with compatible professional capabilities, the routing node serves as the recognition entity; otherwise, the routing node filters out the smart home nodes whose second load conditions meet the second preset condition in the load matrix, and preferentially selects the smart home nodes with special processing advantages for the data type involved in the original parameter as the recognition entity. If there are multiple nodes that meet the conditions, they are sorted in ascending order of the communication delay between the nodes and the routing node, and the smart home node ranked first is selected.

[0015] Further, the first load condition includes the usage rate of the central processing unit, the memory occupancy rate, and the data transmission bandwidth occupancy; the second load condition includes the length of the task processing queue and the proportion of the remaining computing resources; the first preset condition is that the usage rate of the central processing unit is lower than 30%, the memory occupancy rate is lower than 40%, and the data transmission bandwidth occupancy rate is lower than 50%; the second preset condition is that the length of the task processing queue is zero and the proportion of the remaining computing resources is higher than 60%; the data types involved in the original parameter include voice type, image type, and video type.

[0016] Further, the construction and operation method of the preset instruction recognition model are as follows. Let the set of collected original parameters be P = {p1, p2, p3,..., p n}, and the corresponding set of environmental parameters be E = {e1, e2, e3,..., e m}, where n and m are the numbers of the original parameters and the environmental parameters respectively. A command sample library S = {s1, s2, s3,..., s k} is established in advance, k is the number of samples, and each sample s i includes the corresponding standard original parameter and environmental parameter and the standard instruction result

[0017] After obtaining the real-time environmental parameter E and original parameter P, the recognition entity calculates the similarity function where w j is the weight of the jth original parameter, which is dynamically adjusted according to historical data statistics and machine learning algorithms to reflect the importance of different parameters for instruction recognition; at the same time, the environmental similarity w l is the weight of the lth environmental parameter; the instruction matching rate is obtained by combining the two similarities where α and β are preset weight coefficients, and α + β = 1, which are used to balance the influence of the original parameters and the environmental parameters on the matching rate; when MR is greater than or equal to the preset instruction matching threshold T, the standard instruction result corresponding to the sample with the highest matching degree is selected from the sample library S as the instruction recognition result, otherwise it is determined that the instruction matching rate is lower than the preset value, and the subsequent confirmation and assistance recognition process is started.

[0018] Furthermore, the original parameters include voice segments, gesture action information, and touch operation sequences; the preset instruction recognition model uses a pre-trained speech recognition model to convert them into text information. Let the set of converted text information be T = {t1, t2, t3,..., t q}, where q is the number of converted texts;

[0019] Meanwhile, a semantic vector space model is pre-constructed, and a word vector matrix W is trained using a large-scale corpus. For each text t in the text information set T i , it is converted into a semantic vector through the word vector matrix W to obtain a set of semantic vectors

[0020] The corresponding set of environmental parameters is E = {e1, e2, e3,..., e m}, and the environmental parameters are also converted into a set of feature vectors using a pre-trained feature extraction model where m is the number of feature vectors after the conversion of the environmental parameters;

[0021] An instruction sample library is pre-established k is the number of samples, and each sample s i contains the corresponding standard text information and its set of semantic vectors set of standard environmental feature vectors and the standard instruction result

[0022] After the recognition subject obtains the real-time environmental parameters E and the original parameters P, for the text information part, by calculating the text semantic similarity function where w j is the weight of the j-th text semantic vector, which is dynamically adjusted according to historical data statistics and machine learning algorithms to reflect the importance of different texts for instruction recognition; at the same time, the environmental feature vector similarity is calculated w l is the weight of the l-th environmental feature vector; the instruction matching rate is obtained by combining the two similarities Among them, α and β are preset weight coefficients, and α + β = 1, which are used to balance the influence of text information and environmental parameters on the matching rate; when MR is greater than or equal to the preset instruction matching threshold T, the standard instruction result corresponding to the sample with the highest matching degree is selected from the sample library S As the instruction recognition result, otherwise it is determined that the instruction matching rate is lower than the preset value, and the subsequent confirmation and assistance recognition process is started.

[0023] Furthermore, the specific method for the cloud server to construct a regionalized instruction learning database is as follows:

[0024] The received first object information is classified according to geographical location information. It is assumed that the global geographical location is divided into G large regions, each large region is further divided into R small regions, and each small region is uniquely identified by Z ij Indicated, where i = 1, 2,..., G, j = 1, 2,..., R; for each small region ij , a corresponding sub-database D is established ij , and its database structure includes an original parameter table P ij , an environmental parameter table E ij , an instruction recognition result table I ij , and an instruction matching rate table MR ij , where the original parameter table P ij Stores the set of original parameters from the corresponding region k represents the data record serial number, and n is the number of original parameters; the environmental parameter table m stores the set of environmental parameters m is the number of environmental parameters; the instruction recognition result table I ij Stores the set of instruction recognition results p is the number of instruction results; the instruction matching rate table MR ij Stores the set of instruction matching rates q is the number of instruction matching rates.

[0025] Furthermore, in the process of constructing the database, a time series analysis model is introduced. Let the time series be T = {t1, t2, t3,..., t s}}, s is the number of time points. For each small region Z ij , the instruction trend function is calculated based on historical data Among them, L is the number of trigonometric function terms, a ijl , w ijl , and b ijl Are the parameters obtained by fitting the time series data by the least squares method. This function is used to predict the trend change of instructions in this region at different time points; at the same time, the instruction frequency distribution function is calculated Among them, c ijpLet \(n\) be the number of times the instruction result \(p\) appears in this region, and \(P\) be the number of all possible instruction results. This function is used to reflect the distribution frequency of the instruction results in the region.

[0026] Further, when receiving an assistance recognition request, the cloud server first locates the corresponding small region \(Z\) according to the geographical location information in the request ij , and then combines the instruction matching rate \(MR\) and uses the following formula to quickly screen out the most matching historical data:

[0027]

[0028] where is the similarity function between the request instruction matching rate and the historical instruction matching rate, which can be calculated using cosine similarity. \(\alpha\), \(\beta\), \(\gamma\) are pre-set weight coefficients, and \(\alpha+\beta+\gamma = 1\). \(\lambda\) is a adjustment coefficient used to adjust the influence of the time trend on the screening result; according to sort the historical data, and select several pieces of data with the highest scores as the result of quick matching and feedback them to the routing node.

[0029] Further, after the cloud server receives the assistance recognition request, the following quick retrieval operation is performed:

[0030] First, perform feature extraction on the geographical location information \(Loc\) in the assistance request. Divide the earth's surface into \(M\times N\) grid regions according to longitude and latitude, and assign a unique code \(C\) to each region mn , where \(m = 1, 2,\cdots, M\), \(n = 1, 2,\cdots, N\). Determine the grid region code \(C\) corresponding to the request through the positioning algorithm mn , and convert it into a \(K\)-dimensional geographical location feature vector \(V\) Loc =\( vloc1 , v loc2 ,\cdots, v locK \), and the conversion formula is where \(w ij \) is the weight coefficient, and \(f(C mn , i)\) is a feature function based on the grid region code \(C mn \) and the dimension index \(i\). \(I\) is the number of intermediate calculation parameters;

[0031] At the same time, standardize the instruction matching rate \(MR\). Let the received instruction matching rate be \(MR req , and convert it into a \(Z\)-score under the standard normal distribution. The calculation formula is where \(\mu\) is the mean of the historical instruction matching rate, and \(\sigma\) is the standard deviation, and obtain the standardized instruction matching rate eigenvalue \(Z MR ;

[0032] Then, construct a retrieval matching model, and pre-establish an instruction result database \(DB\) result, where each record contains a geographical location feature vector Normalized instruction matching rate eigenvalue And the corresponding instruction recognition result r is the record index; the cosine similarity function is used to calculate the geographical location similarity between the request and the database record, and the Gaussian kernel function is used to calculate the instruction matching rate similarity, where δ is the Gaussian kernel bandwidth;

[0033] Finally, combining the two similarities, through the weighted summation formula:

[0034] Calculate the matching score Score for each record r , where α and β are preset weight coefficients, and α + β = 1; sort the database records according to the matching score, and select the instruction recognition results corresponding to the top P records as the fast matching results and feedback them to the routing node.

[0035] The scheme principle and beneficial effects are as follows: Multiple smart home nodes have diverse data collection capabilities. They are distributed throughout the home and can real-time sense environmental parameters such as light, temperature, humidity, and noise. At the same time, they can accurately capture the original parameters input by users, which cover various modal information such as voice commands, gesture actions, and touch operations. For example, when the user says "turn on the TV" to the smart speaker in the living room, this is the voice modality; if the user is busy in the kitchen with both hands covered in stains and points to the direction of the light through a simple waving gesture, the camera of the smart home node captures this gesture action and recognizes it as a light-on command, this is the gesture modality; another example is adjusting the air conditioner temperature by touching the smart control panel in the bedroom, which belongs to the touch modality.

[0036] The multi-modal original parameters and environmental parameters collected are sent to the routing node. The routing node determines the instruction recognition subject according to the preset allocation rules, taking into account its own and the load conditions of each smart home node. If the routing node has a low load and its built-in voice recognition module can efficiently process voice commands, the routing node will recognize the voice commands; if it encounters image-like gesture data and a certain smart home node is equipped with a professional image recognition chip and is idle, the gesture data will be allocated to this node for recognition.

[0037] The original parameters collected by the smart home node are attached with geographical location information, and this information and the instruction data are sent by the routing node to the cloud server. The cloud server constructs a regionalized instruction learning database with geographical location information and instruction matching rate as the key indexes. For example, the world is divided into different large regions and small regions, and a corresponding sub-database is established for each small region to store data such as local user instruction habits, environmental characteristics, and instruction results.

[0038] The cloud server quickly locates the corresponding regional database based on the geographical location information in the instruction. Combining the instruction matching rate, it uses a specific algorithm to filter out the historical instruction data most similar to the current instruction, achieving precise adaptation to regional instruction differences. For example, in the northern region, users often say "Turn up the heating", while in the southern region, it may be "Turn up the heating temperature of the air conditioner". The system improves the accuracy of the first instruction recognition by learning the instruction expression methods in different regions.

[0039] In the instruction recognition stage, whether it is the routing node itself or the selected smart home node as the recognition subject, it will obtain the environmental parameters and original parameters and process them according to the preset instruction recognition model. This model comprehensively considers various information such as the text semantics of the original parameters and the feature vectors of the environmental parameters. For example, through the pre-constructed semantic vector space model, the text information after voice conversion is converted into a semantic vector, and at the same time, the environmental parameters are converted into feature vectors.

[0040] Calculate the similarity of text semantics and the similarity of environmental feature vectors, and comprehensively obtain the instruction matching rate according to the preset weight coefficient. If the instruction matching rate is lower than the preset value, actively send an instruction confirmation request to the user terminal and send an assistance recognition request to the cloud server; if it is higher than the preset value, directly determine the instruction recognition result and drive the corresponding smart home device to execute the operation, realizing intelligent decision-making combining multiple key information.

[0041] Multi-modal interaction integration enables users to conveniently control smart home devices in different scenarios. Whether it is easy voice control when relaxing and watching TV, gesture operation when busy with housework and having inconvenient hands, or touch control when not wanting to make a sound at night, it can meet the needs, greatly improving the flexibility and convenience of interaction and adapting to various life scenarios.

[0042] After introducing the logic of learning instruction recognition from adjacent or nearby families, the intelligent learning boundary of the system is further expanded. When users input some relatively novel or personalized instructions, the system is not only limited to the data accumulation of its own family in the past, but also can refer to the instruction patterns of adjacent families. For example, newly married young couples may use some popular and creative instructions to control smart homes, such as "Turn on the romantic dinner mode". If there are similar settings in the surrounding neighbors, the system can quickly learn and adapt, providing users with a more trendy and personalized interaction experience, making users feel the progressiveness of smart homes.

[0043] Precisely adapting to regional instruction differences enables the system to understand instructions according to the language habits and cultural customs of the user's location. There will no longer be situations where users in different regions have incorrect instruction recognition or repeated confirmations due to a unified instruction vocabulary, reducing the complexity of user operations, improving the system response speed, and making smart home interactions more fluent.

[0044] With the help of the instruction data of adjacent or nearby households, the limitations of individual household instruction samples can be effectively filled. Some niche but regionally characteristic instructions may occur very rarely in single-household data and it is difficult to form an effective recognition pattern. However, by learning from surrounding households, the instruction samples can be quickly enriched, improving the recognition accuracy of such special instructions. For example, in a certain characteristic old street, local residents are used to using specific dialect words to control smart door locks. Through learning among neighbors, the system can accurately recognize these unique instructions and avoid misjudgment.

[0045] By combining multiple key information for intelligent decision-making, the system can adjust the instruction recognition strategy in real time according to factors such as environmental changes and user behavior habits. In the early morning, the curtains are automatically opened and soothing music is played according to the brightening of the light; during a family gathering, the system automatically creates a movie-watching atmosphere based on the keywords of multi-person communication, making the operation of smart home devices more in line with the actual needs of users, truly realizing an intelligent life and improving the quality of home life.

[0046] Integrating the neighborhood instruction learning mechanism makes the system decision-making more intelligent with group wisdom. When dealing with emergencies within the community, such as a rainstorm warning, if some neighbors take the lead in activating the "rainstorm protection mode" (closing windows, retracting outdoor drying racks, etc.), the system can recommend this mode to surrounding households based on the principle of proximity in geographical location, realizing group intelligent linkage, preventing risks in advance, and providing more comprehensive protection for users' lives. Brief Description of the Drawings

[0047] Figure 1 It is the front view / cross-sectional view of the first embodiment of the smart home intelligent interaction system and method based on big data. Detailed Implementation Modes

[0048] The following is a further detailed description through specific implementation modes:

[0049] The smart home intelligent interaction system based on big data includes: a cloud server, a user terminal, a routing node, and several smart home nodes.

[0050] Cloud server: Select a cloud service device with high computing power and large-capacity storage, such as the Amazon AWS series of servers, and configure a high-performance CPU, a large amount of memory, and a solid-state drive with fast read and write speeds. Build a data management platform specifically for the smart home interaction system on the cloud server, and install database management software, such as MySQL or MongoDB, for building and maintaining a regional instruction learning database.

[0051] User terminal: It can be a smart phone or a tablet computer. Taking a smart phone as an example, it installs a customized smart home control APP and communicates with the routing node through Wi-Fi or 4G / 5G network. The APP interface is designed to be simple and intuitive, with a voice input button, a gesture guidance tutorial entry, and a quick operation area for common instructions, facilitating users to quickly issue instructions. At the same time, the smart phone enables the permissions of the built-in gyroscope and acceleration sensor to continuously monitor the motion state of the phone and provide negative feedback data for instruction learning.

[0052] Routing node: Select a wireless router that supports multi-device connection and has an intelligent load balancing function, such as the Huawei AX3 Pro. It is deployed at the center of the home to ensure good Wi-Fi signal coverage throughout the house. A customized intelligent control firmware is burned on the routing node to enable it to stably communicate with the cloud server, smart home nodes, and user terminals. At the same time, a preliminary instruction recognition module is built in to quickly preprocess simple voice and text instructions.

[0053] The smart home nodes are distributed in the home environment and are used to collect environmental parameters and raw parameters input by users in real time, and send the environmental parameters and raw parameters to the routing node. The environmental parameters include geographical location information.

[0054] Specific smart home nodes include environmental perception nodes and interaction execution nodes. The environmental perception nodes are distributed in each room, such as the living room, bedroom, kitchen, bathroom, etc. A multi-functional sensor integration module is used, such as the Xiaomi Mijia temperature and humidity sensor, light sensor, etc., to collect environmental parameters such as light intensity, temperature, humidity, and noise level in real time, update the data every 5 seconds, and transmit the data to the routing node through the ZigBee or Bluetooth Low Energy protocol.

[0055] Interaction execution nodes such as smart speakers (such as the Xiaoai Tongxue speaker) are used to receive voice instructions and play feedback information. Smart bulbs (such as the Philips Hue smart light) can adjust the brightness and color according to instructions. Smart curtain motors (such as the Dooya smart curtain motor) control the opening and closing of the curtains. These nodes are built with a microprocessor and have a certain instruction parsing ability, which can directly execute simple instructions and can also forward complex instructions to the routing node for processing.

[0056] The routing node is used to determine the routing node or a certain smart home node as the recognition subject for this instruction recognition after receiving the original parameters according to the preset allocation rule; the recognition subject obtains the environmental parameters and the original parameters and obtains the instruction recognition result and the instruction matching rate according to the preset instruction recognition model; if the instruction matching rate is lower than the preset value, an instruction confirmation request is actively sent to the user terminal, and at the same time, an assistance recognition request is sent to the cloud server. If the instruction confirmation information or the modified instruction recognition result input by the user through the user terminal is received within the preset time, the instruction confirmation information or the modified instruction recognition result is sent to the corresponding smart home node; then the environmental parameters, the original parameters, and the instruction recognition result in this recognition are combined into the first object information and sent to the cloud server.

[0057] The cloud server is used for the first object information, and constructs a regionalized instruction learning database with the geographical location information and the instruction matching rate therein as indexes; and after receiving the assistance recognition request, retrieves with the geographical location information and the instruction matching rate corresponding to the assistance request as indexes, performs rapid matching of the instruction recognition result, and then feeds back the successfully matched instruction recognition result to the routing node.

[0058] Among them, if the routing node does not receive the user feedback within the preset time, the instruction recognition result sent by the cloud server received is directly sent to the corresponding smart home node.

[0059] In specific use: When the user says "Dim the living room lights a bit" to the Xiaoai speaker in the living room, the voice instruction is captured by the smart speaker. The speaker first performs preliminary noise reduction and audio feature extraction processing on the voice, and then sends a data packet containing the original audio data of the voice instruction, its own device ID (used to identify the location information, such as device No. 01 in the living room), and the currently collected living room environmental parameters (light intensity XX lux, temperature XX °C, humidity XX%, noise XX dB) to the routing node through Wi-Fi.

[0060] If the user waves to indicate turning on the range hood in the kitchen, the smart camera installed on the kitchen ceiling (as one of the smart home nodes) captures this gesture action, uses the built-in image recognition algorithm to initially judge the meaning of the gesture, and packs the gesture action data, the location information of the camera (device No. 02 in the kitchen), and the real-time environmental parameters of the kitchen, and sends them to the routing node through wired Ethernet or Wi-Fi.

[0061] After receiving a data packet, the routing node immediately activates the load balancing strategy. It first checks its own CPU usage rate, memory occupancy rate, and the number of instructions being processed currently. Suppose the CPU usage rate of the routing node has exceeded 70%. At this time, it sends a load query instruction to all the smart home nodes in the whole house. Each node feeds back information about the length of its task processing queue and the proportion of remaining computing resources within 0.5 seconds. The routing node finds that the smart camera in the kitchen has abundant remaining computing resources, and its image recognition ability is suitable for gesture instruction processing. So it forwards the gesture instruction data packet sent from the kitchen to this camera for further processing.

[0062] As the recognition subject, after obtaining the complete environmental parameters and the original gesture action parameters, the smart camera processes them according to the preset instruction recognition model. A small gesture instruction sample library is pre-stored locally in the camera. Let the set of collected gesture action parameters be G = {g1, g2, g3, …, g n}, the corresponding set of environmental parameters be E = {e1, e2, e3, …, e m}, the sample library S = {s1, s2, s3, ..., s k}, and each sample s i contains the corresponding standard gesture parameters and environmental parameters as well as the standard instruction results By calculating the similarity function where w j is the weight of the jth gesture parameter, which is dynamically adjusted according to historical data statistics and machine learning algorithms. At the same time, calculate the environmental similarity w l is the weight of the lth environmental parameter; combining the two similarities to obtain the instruction matching rate where α and β are preset weight coefficients, and α + β = 1. Suppose the calculated instruction matching rate MR is lower than the preset value of 60%. The smart camera immediately sends an instruction confirmation request to the smart home control APP on the user's mobile phone through the routing node, and at the same time sends an assistance recognition request to the cloud server.

[0063] After receiving the assistance recognition request, the cloud server locates the corresponding small area database according to the geographical location information (the area code corresponding to device No. 02 in the kitchen) in the request. Suppose this database has stored the instruction data of 100 surrounding households in the kitchen scenario. The cloud server uses the following formula to quickly screen out the most matching historical data:

[0064]

[0065] where is the similarity function between the request instruction matching rate and the historical instruction matching rate, which can be calculated by cosine similarity. α, β, and γ are preset weight coefficients, and α + β + γ = 1. λ is an adjustment coefficient used to adjust the influence of the time trend on the screening result; F ij (t) is the instruction trend function calculated based on historical data, H ij (p) is the instruction frequency distribution function. According to Sort the historical data, and select the 5 data with the highest scores as the result of quick matching and feedback it to the routing node.

[0066] If the routing node does not receive user feedback within the preset time (such as 10 seconds), it directly sends the instruction recognition result sent by the cloud server to the corresponding smart home node. Assuming that the instruction recognition result feedback by the cloud server is "turn on the kitchen range hood", the routing node will send this instruction to the smart range hood in the kitchen, and the range hood will perform the corresponding action to complete a complete instruction interaction process. At the same time, the routing node combines the environmental parameters, original parameters, and instruction recognition result in this recognition into the first object information and sends it to the cloud server, and the cloud server updates the corresponding area database to provide more accurate data support for subsequent instruction recognition.

[0067] The cloud server automatically scans the instruction data update situation of families with similar geographical locations (such as the same building and the same unit) in the same community once a week. When it is found that neighbor A's family has newly used an instruction "activate the kitchen air purification mode" and this instruction has not appeared in neighboring families, the cloud server includes it in a temporary learning pool.

[0068] For the instructions included in the learning pool, the cloud server preferentially analyzes information such as the environmental parameters and device operation logic corresponding to the instructions. Assuming that when neighbor A's family activates the kitchen air purification mode, the environmental parameters show that the kitchen smoke concentration is relatively high and the temperature is slightly high. The cloud server combines this information. In the next week, if the similar environmental parameters appear in its own kitchen and the user input instruction is unclear, it actively pushes the suggestion of "whether to activate the kitchen air purification mode" to the user terminal. In this way, it gradually learns and integrates into the advanced instruction modes of neighboring families, continuously improving the intelligence level of the system.

[0069] In other embodiments, a large amount of speech segment data is collected, covering different accents, speech rates, tones, etc. This data is sourced from the usage records of actual users, publicly available speech datasets, etc. The speech data is labeled with corresponding text information to form training samples. A deep learning architecture, such as a convolutional neural network (CNN) combined with a recurrent neural network (RNN) or its variants (such as LSTM, GRU), is used to build a speech recognition model. During training, stochastic gradient descent (SGD) and its optimization algorithms (such as Adam, Adagrad) are used to iteratively update the model parameters, continuously adjusting the model to minimize the error between the predicted text and the labeled text, such as the cross-entropy loss function. After multiple rounds of training, a speech recognition model that can accurately convert speech segments into text information is obtained.

[0070] A large-scale corpus is prepared, which can be news articles, novels, professional documents, etc. Algorithms such as Word2Vec, GloVe, or FastText are used to train the corpus. Taking Word2Vec as an example, by constructing a context window of words, the target word is predicted, thereby learning the vector representation of each word. During the training process, the word vector matrix is continuously adjusted so that words with similar semantics are closer in the vector space. Finally, a word vector matrix that can convert text information into semantic vectors is obtained.

[0071] A large amount of environmental parameter data is collected, including light intensity, temperature, humidity, noise level, etc. Feature engineering is performed on this data, such as normalization and standardization. Machine learning algorithms, such as support vector machines (SVM), decision trees, or multi-layer perceptrons (MLP) in deep learning, are used to build an environmental feature extraction model. The environmental parameters are used as input to extract vectors that can represent environmental features. The model parameters are continuously adjusted through training data so that the model can accurately convert environmental parameters into feature vectors.

[0072] Smart home nodes are distributed at various locations in the home environment, such as the living room, bedroom, kitchen, etc. Each node is equipped with corresponding sensors and interaction devices, such as a microphone for collecting speech segments, a camera for capturing gesture action information, a touch screen for recording touch operation sequences, and environmental sensors are also installed to obtain environmental parameters such as light intensity, temperature, humidity, and noise level.

[0073] When the user issues an instruction, for example, saying "turn on the TV" to the smart speaker in the living room, the microphone collects the speech segment, and at the same time, the environmental sensors record the current environmental parameters such as light intensity and temperature in the living room. The smart home node packs the original parameters (speech segment) and environmental parameters and sends them to the routing node through wireless communication protocols such as ZigBee and Wi-Fi.

[0074] After receiving the original parameters and environmental parameters, the routing node determines the instruction recognition entity according to the preset allocation rule. The preset allocation rule can be based on the load conditions of the routing node and each smart home node. For example, if the CPU usage rate of the routing node is lower than 30% and the memory occupancy rate is lower than 40%, then the routing node serves as the recognition entity; otherwise, a smart home node with a lighter load and corresponding processing capabilities is selected as the recognition entity.

[0075] After obtaining the environmental parameters and original parameters, the recognition entity processes them according to the preset instruction recognition model. First, the speech recognition model trained in advance is used to convert the speech segment into text information. Suppose the set of converted text information is T = {t1, t2, t3, …, t q}, and q is the number of converted texts.

[0076] Then, through the pre-constructed semantic vector space model, each text t in the text information set T is converted into a semantic vector by using the word vector matrix W i to obtain the semantic vector set

[0077] Meanwhile, the corresponding set of environmental parameters is E = {e1, e2, e3, …, e m}, and the environmental parameters are also converted into a set of feature vectors by using the pre-trained feature extraction model where m is the number of feature vectors after the conversion of environmental parameters.

[0078] An instruction sample library S = {s1, s2, s3, …, s k} is established in advance, k is the number of samples, and each sample s i contains the corresponding standard text information and its semantic vector set standard environmental feature vector set and standard instruction result

[0079] For the text information part, the recognition entity calculates the text semantic similarity function w j is the weight of the j-th text semantic vector, which is dynamically adjusted according to historical data statistics and machine learning algorithms to reflect the importance of different texts for instruction recognition. At the same time, the environmental feature vector similarity where w l is the weight of the l-th environmental feature vector.

[0080] The instruction matching rate is obtained by combining the two similarities Among them, α and β are preset weight coefficients, and α + β = 1, which are used to balance the influence of text information and environmental parameters on the matching rate; when MR is greater than or equal to the preset instruction matching threshold T, the standard instruction result corresponding to the sample with the highest matching degree is selected from the sample library S As the instruction recognition result, otherwise it is determined that the instruction matching rate is lower than the preset value, and the subsequent confirmation and assistance recognition process is started.

[0081] The cloud server classifies the received first object information according to the geographical location information. Suppose the global geographical location is divided into G large regions, each large region is further divided into R small regions, and each small region is uniquely identified by Z ij Indicates, where i = 1, 2,..., G, j = 1, 2,..., R; for each small region Z ij , a corresponding sub-database D ij is established, and its database structure includes the original parameter table P ij , the environmental parameter table E ij , the instruction recognition result table I ij and the instruction matching rate table MR ij , where the original parameter table P ij stores the original parameter set from the corresponding region k represents the data record serial number, and n is the number of original parameters; the environmental parameter table E ij stores the environmental parameter set m is the number of environmental parameters; the instruction recognition result table I ij stores the instruction recognition result set p is the number of instruction results; the instruction matching rate table MR ij stores the instruction matching rate set q is the number of instruction matching rates.

[0082] When the cloud server receives an assistance recognition request, it first performs feature extraction on the geographical location information Loc in the assistance request. The earth's surface is divided into M×N grid regions according to longitude and latitude, and each region is assigned a unique code C mn , where m = 1, 2,..., M, n = 1, 2,..., N. The grid region code C mn corresponding to the request is determined through the positioning algorithm and converted into a K-dimensional geographical location feature vector V Loc = [v loc1 , v loc2 , …, v locK , and the conversion formula is where w ij is the weight coefficient, f(C mn , i) is the feature function based on the grid region code C mn and the dimension index i, and I is the number of intermediate calculation parameters;

[0083] Meanwhile, standardize the instruction matching rate MR. Let the received instruction matching rate be MR req , and convert it to the Z-score under the standard normal distribution. The calculation formula is where μ is the mean of the historical instruction matching rate, and σ is the standard deviation, obtaining the standardized instruction matching rate eigenvalue Z MR ;

[0084] Then, construct a retrieval matching model and pre-establish an instruction result database DB result , where each record contains a geographical location feature vector standardized instruction matching rate eigenvalue and the corresponding instruction recognition result r is the record index; use the cosine similarity function to calculate the geographical location similarity between the request and the database record, and use the Gaussian kernel function to calculate the instruction matching rate similarity, where δ is the Gaussian kernel bandwidth;

[0085] Finally, combine the two similarities through the weighted summation formula:

[0086] Calculate the matching score Score for each record r , where α and β are preset weight coefficients, and α + β = 1; sort the database records according to the matching score, and select the instruction recognition results corresponding to the top P records as the fast matching results and feedback them to the routing node.

[0087] Through multi-stage model training, the speech recognition model can adapt to the speech characteristics of different users and accurately convert speech segments into text information. The introduction of the semantic vector space model and the environmental feature extraction model enables the comprehensive consideration of text semantics and environmental factors during the instruction recognition process. When calculating the instruction matching rate, by combining the text semantic similarity and the environmental feature vector similarity, it is possible to more accurately judge the matching degree between the user's instruction and the standard instruction in the sample library. Meanwhile, the regionalized instruction learning database constructed by the cloud server classifies the data according to geographical location information, enabling the system to better adapt to the instruction habits and environmental characteristics of users in different regions, further improving the accuracy of instruction recognition. For example, in the humid southern regions, users may use the instruction "turn on the dehumidification mode" more frequently. By learning the data in this region, the system can more accurately recognize and respond to such instructions.

[0088] The preset allocation rule of the routing node can quickly determine the instruction recognition entity according to the load situation, avoiding unnecessary waiting and resource waste. After receiving the assistance recognition request, the cloud server can, through the characterization processing of the geographical location information and the instruction matching rate, as well as the fast retrieval algorithm, find the matching instruction recognition result from the huge instruction result database within a short time and feedback it to the routing node. For example, when the user issues a complex instruction and the local recognition matching rate is low, the cloud server can complete the retrieval and feedback within milliseconds, greatly shortening the user's waiting time for the response and improving the user experience. This fast response ability enables the smart home system to more timely meet the user's needs and allows the user to experience an efficient and convenient smart life.

[0089] The above are only the embodiments of the present invention. The present invention is not limited to the fields involved in this embodiment. Common knowledge such as the specific structures and characteristics known in the art is not described in detail herein. Those of ordinary skill in the art know all the common technical knowledge in the technical field to which the invention belongs before the application date or the priority date, can know all the existing technologies in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to improve and implement this solution. Some typical known structures or known methods should not become obstacles for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners described in the specification can be used to explain the content of the claims.

Claims

1. The smart home intelligent interactive system based on big data is characterized by: include: Cloud servers, user terminals, routing nodes, and several smart home nodes; The smart home nodes are distributed in the home environment, and are used to collect environmental parameters and original parameters input by users in real time, and send the environmental parameters and original parameters to the routing node, wherein the environmental parameters include geographical location information; The routing node is used to determine the routing node or a smart home node as the identification subject of this command identification according to the preset allocation rule after receiving the original parameters; the identification subject obtains the environmental parameters and the original parameters to obtain the command identification result and the command matching rate according to the preset command identification model; If the command matching rate is lower than the preset value, a command confirmation request is actively sent to the user terminal, and an assist recognition request is sent to the cloud server at the same time. If the command confirmation information or the modified command recognition result input by the user through the user terminal is received within the preset time, the command confirmation information or the modified command recognition result is sent to the corresponding smart home node; Then the environmental parameters, original parameters and instruction recognition results in this recognition are combined into the first object information and sent to the cloud server; The cloud server is used for the first object information, and uses the geographical location information and the command matching rate therein as an index to build a regionalized command learning database; and after receiving an assistance recognition request, uses the geographical location information and the command matching rate corresponding to the assistance request as an index to perform a search, perform a fast match of the command recognition result, and then feeds back the successfully matched command recognition result to the routing node; If the routing node does not receive user feedback within a preset time, the command recognition result sent by the cloud server will be directly sent to the corresponding smart home node.

2. The smart home intelligent interaction system based on big data according to claim 1 is characterized in that The preset allocation rule is that the routing node monitors its own first load condition in real time at a frequency of at least once per second, and sends a load query instruction to each smart home node, and each smart home node feeds back its own second load condition after receiving the instruction; the routing node builds a load matrix according to the first load condition and the second load condition; when the original parameter arrives, if the load of the routing node meets the first preset condition, and there is no idle smart home node with professional capability adaptation, the routing node serves as the identification subject; otherwise, the routing node selects the smart home node whose second load condition meets the second preset condition in the load matrix, and gives priority to the smart home node with special processing advantages for the data type involved in the original parameter as the identification subject. If there are multiple nodes that meet the conditions, they are sorted from small to large according to the communication delay between the node and the routing node, and the smart home node ranked first is selected.

3. The smart home intelligent interactive system based on big data according to claim 2 is characterized in that ,, the first load condition includes the CPU usage rate, memory occupancy rate and data transmission bandwidth occupancy; the second load condition includes the task processing queue length and the remaining computing resource ratio information; the first preset condition is that the CPU usage rate is lower than 30%, the memory occupancy rate is lower than 40% and the data transmission bandwidth occupancy rate is lower than 50%; the second preset condition is that the task processing queue length is zero and the remaining computing resource ratio is higher than 60%; the data types involved in the original parameters include voice type, image type and video type.

4. The smart home intelligent interactive system based on big data according to claim 3 is characterized in that: The preset instruction recognition model is constructed and operated as follows: Assume that the collected original parameter set is P = {p1, p2, p3, ..., p n }, the corresponding environmental parameter set is E = {e1, e2, e3, ..., e m }, where n and m are the number of original parameters and environmental parameters respectively; a pre-established instruction sample library S = {s1, s2, s3, ..., s k }, k is the number of samples, each sample s i Contains the original parameters of the corresponding standard and environmental parameters And the standard instruction results After obtaining the real-time environmental parameters E and original parameters P, the recognition subject calculates the similarity function where w j is the weight of the jth original parameter, which is dynamically adjusted based on historical data statistics and machine learning algorithms to reflect the importance of different parameters to command recognition; at the same time, the environmental similarity is calculated w l is the weight of the lth environmental parameter; Combining the two similarities to get the instruction matching rate Where α and β are pre-set weight coefficients, and α+β=1, which is used to balance the influence of original parameters and environmental parameters on the matching rate; when MR is greater than or equal to the preset instruction matching threshold T, the standard instruction result corresponding to the sample with the highest matching degree is selected from the sample library S As the result of instruction recognition, otherwise it is determined that the instruction matching rate is lower than the preset value, and the subsequent confirmation and auxiliary recognition process is started.

5. The smart home intelligent interactive system based on big data according to claim 3 is characterized in that: The original parameters include speech segments, gesture action information and touch operation sequences; the preset command recognition model converts them into text information using a pre-trained speech recognition model, and the converted text information set is T = {t1, t2, t3, ..., t q }, q is the number of converted texts; At the same time, a semantic vector space model is pre-built, and a word vector matrix W is obtained by training with a large-scale corpus. For each text t in the text information set T, i , which is converted into a semantic vector through the word vector matrix W Get semantic vector set The corresponding environmental parameter set is E = {e1, e2, e3, ..., e m }, also using the pre-trained feature extraction model to convert environmental parameters into a set of feature vectors Where m is the number of eigenvectors after environmental parameter conversion; Pre-establish the instruction sample library S = {s1, s2, s3, ..., s k }, k is the number of samples, each sample s i Contains the corresponding standard text information and its semantic vector set Standard environment feature vector set And the standard instruction results After the recognition subject obtains the real-time environmental parameters E and original parameters P, for the text information part, it calculates the text semantic similarity function where w j is the weight of the jth text semantic vector, which is dynamically adjusted based on historical data statistics and machine learning algorithms to reflect the importance of different texts to command recognition; at the same time, the similarity of the environmental feature vector is calculated w l is the weight of the lth environment feature vector; the instruction matching rate is obtained by combining the two similarities Where α and β are pre-set weight coefficients, and α+β=1, which is used to balance the influence of text information and environmental parameters on the matching rate; when MR is greater than or equal to the preset instruction matching threshold T, the standard instruction result corresponding to the sample with the highest matching degree is selected from the sample library S As the result of instruction recognition, otherwise it is determined that the instruction matching rate is lower than the preset value, and the subsequent confirmation and auxiliary recognition process is started.

6. The smart home intelligent interactive system based on big data according to claim 5 is characterized in that: The specific method of the cloud server to construct a regionalized instruction learning database is as follows: The received first object information is classified according to the geographical location information. Assume that the global geographical location is divided into G regions, each of which is further divided into R small regions, and each small region is uniquely identified by Z. ij Represents, where i = 1, 2, ..., G, j = 1, 2, ..., R; for each small area Z ij , establish the corresponding sub-database D ij , whose database structure includes the original parameter table P ij 、Environmental parameter table E ij , Instruction recognition results Table I ij And the instruction matching rate table MR ij , where the original parameter table P ij Stores the original parameter set from the corresponding region k represents the data record number, n is the number of original parameters; Environmental parameter table E ij Storage environment parameter collection m is the number of environmental parameters; instruction recognition results table I ij Store instruction recognition result set p is the number of instruction results; instruction matching rate table MR ij Storage instruction match rate set q is the number of instruction matching rates.

7. The smart home intelligent interactive system based on big data according to claim 6 is characterized in that: In the process of building the database, a time series analysis model is introduced. Assume that the time series is T = {t1, t2, t3, ..., t s }, s is the number of time points, for each small area Z ij , calculate instruction trend function based on historical data Where L is the number of trigonometric functions, a ijl ,ω ijl , and b ijl is the parameter obtained by fitting the time series data by the least squares method. This function is used to predict the trend change of instructions in the area at different time points. At the same time, the instruction frequency distribution function is calculated. where c ijp is the number of times the instruction result p appears in the area, P is the number of all possible instruction results, and this function is used to reflect the distribution frequency of instruction results in the area.

8. The smart home intelligent interaction system based on big data according to claim 7 is characterized in that: When receiving an assistance recognition request, the cloud server first locates the corresponding small area Z according to the geographic location information in the request. ij , and then combined with the instruction matching rate MR, use the following formula to quickly filter out the most matching historical data: in is the similarity function between the request instruction matching rate and the historical instruction matching rate, which can be calculated by cosine similarity. α, β, and γ are pre-set weight coefficients, and α+β+γ=1. λ is the adjustment coefficient used to adjust the influence of time trend on the screening results. Sort the historical data, select the data with the highest scores as the quick matching results and feed them back to the routing node.

9. The smart home intelligent interactive system based on big data according to claim 8, characterized in that: After the cloud server receives the assistance identification request, it performs the following quick retrieval operations: First, the geographic location information Loc in the assistance request is characterized, and the earth's surface is divided into M×N grid areas according to longitude and latitude, and each area is given a unique code C mn , where m = 1, 2, ..., M, n = 1, 2, ..., N, the grid area code C corresponding to the request is determined by the positioning algorithm mn , and convert it into a K-dimensional geographic location feature vector V Loc =[v loc1 , v loc2 , …, v locK ], the conversion formula is where w ij is the weight coefficient, f(C mn ,i) is based on the grid area coding C mn and the characteristic function of dimension index i, where I is the number of intermediate calculation parameters; At the same time, the instruction matching rate MR is standardized, and the received instruction matching rate is MR req , convert it into the Z score under the standard normal distribution, the calculation formula is Where μ is the mean of the historical instruction matching rate, σ is the standard deviation, and the standardized instruction matching rate characteristic value Z is obtained MR ; Then, build a search matching model and pre-establish the instruction result database DB result , where each record contains a geographic location feature vector Standardized instruction matching rate characteristic value And the corresponding instruction recognition results r is the record index; the cosine similarity function is used to calculate the geographic location similarity between the request and the database record, and the Gaussian kernel function is used Calculate the instruction matching rate similarity, where δ is the Gaussian kernel bandwidth; Finally, the two similarities are combined through the weighted summation formula: Calculate the matching score of each record r , where α and β are pre-set weight coefficients, and α+β=1; the database records are sorted according to the matching scores, and the instruction recognition results corresponding to the first P records are selected as the fast matching results and fed back to the routing node.

10. A smart home intelligent interaction method based on big data, characterized in that: A system as claimed in any one of claims 1 to 9 is employed.