Voice control method and system for smart home

By constructing a smart home network, collecting and analyzing the location and status characteristics of human voice signals, and rationally allocating voice recognition tasks, the problem of misjudgment and missed judgment in voice control in smart home environments has been solved, achieving accurate voice control and device management.

CN121125380AInactive Publication Date: 2025-12-12SHENZHEN JIAYIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511602314.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2025-12-12
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN121125380A_ABST
    Figure CN121125380A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart home, in particular to a voice control method and system for smart home, and the method comprises the steps: carrying out the local area network connection of each piece of smart furniture through a network module disposed in each piece of smart furniture, so as to construct a smart home network; the method comprises the following steps: acquiring a sound signal based on a smart home network, analyzing to obtain a position feature and a state feature of a human voice source, allocating a voice recognition task to specified smart furniture in the smart home network according to the position feature and the state feature, and performing feature recognition on the acquired sound signal based on the voice recognition task. And generating a voice control instruction and sending the voice control instruction to the corresponding intelligent furniture through the intelligent home network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of smart home, more particularly, to a voice control method and system for smart home. BACKGROUND

[0002] In a smart home environment, there will be various forms of environmental noise, which will interfere with the voice control of smart furniture. Meanwhile, in the smart home environment, users will also communicate with each other and use electronic products such as televisions and mobile phones. In these scenarios, misjudgment and missed judgment of smart furniture are likely to occur. SUMMARY

[0003] Therefore, embodiments of the present application provide a voice control method and system for smart home to accurately obtain voice control of smart home.

[0004] To achieve the above object, embodiments of the present application provide the following technical solutions. According to one aspect of the present application, a voice control method for smart home is provided, comprising: connecting each piece of smart furniture through a network module arranged in each piece of smart furniture to build a smart home network; collecting sound signals based on the smart home network and analyzing to obtain position characteristics and state characteristics of a human voice source; allocating a voice recognition task for a designated smart furniture in the smart home network according to the position characteristics and the state characteristics; performing feature recognition on the collected sound signals based on the voice recognition task, generating a voice control instruction and sending it to the corresponding smart furniture through the smart home network.

[0005] According to another aspect of the present application, a voice control system for smart home is provided, comprising: a network connection module for connecting each piece of smart furniture through a network module arranged in each piece of smart furniture to build a smart home network; a feature analysis module for collecting sound signals based on the smart home network and analyzing to obtain position characteristics and state characteristics of a human voice source; a task allocation module for allocating a voice recognition task for a designated smart furniture in the smart home network according to the position characteristics and the state characteristics; a voice control module for performing feature recognition on the collected sound signals based on the voice recognition task, generating a voice control instruction and sending it to the corresponding smart furniture through the smart home network.

[0006] Via the technical solution, the voice control method of the smart home has the following beneficial effects: The application constructs a network to realize device interconnection and intercommunication, analyzes the position and state characteristics of the human voice, accurately locates the sounder, grasps the intention, reasonably allocates the voice recognition task, gives play to the advantages of each smart furniture, improves the recognition efficiency and accuracy, generates accurate control instructions and stably transmits, can realize effective control of the smart furniture, allows the user to conveniently control the home by voice, improves the life comfort and intelligent level, and meets the needs of modern fast-paced life. BRIEF DESCRIPTION OF DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor based on the provided drawings: Figure 1 The step schematic diagram of the voice control method of the smart home provided by the embodiment of the application; Figure 2 The structure schematic diagram of the voice control system of the smart home provided by the embodiment of the application. DETAILED DESCRIPTION

[0008] The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0009] In the smart home environment, there will be various forms of environmental noise, which will interfere with the voice control of the smart furniture. At the same time, in the smart home environment, the user will also communicate with each other and use electronic products such as TV and mobile phone. In these situations, the smart furniture is easy to misjudge and miss.

[0010] Therefore, the application provides a voice control method of smart home, as shown in the steps of Figure 1 , including: First, the network module arranged in each smart furniture is used to locally connect each smart furniture to construct a smart home network.

[0011] Specifically, in the first step of the embodiments provided by the present application, the network module in the smart furniture is turned on to enter a configurable state, which involves powering on the network module, loading a default configuration file, and other operations. Different types of network modules have different initialization methods. For example, a Wi-Fi module needs to perform some basic parameter settings, such as channel and frequency band; a Bluetooth module needs to be turned on to discover mode, etc. The network module is in an unready state before initialization and cannot perform subsequent network connection operations. Initialization can ensure the normal operation of the basic functions of the network module and provide a basis for subsequent network configuration and connection.

[0012] More specifically, the network module of each smart furniture automatically detects the available network environment around it. For modules that support Wi-Fi connection, it will scan the surrounding Wi-Fi hotspots to obtain information such as SSID (network name) and signal strength. For modules that support other wireless protocols such as ZigBee, it will detect the corresponding wireless signal frequency band to see if there is a usable network. Understanding the surrounding network environment is a prerequisite for selecting the appropriate network connection method. Different environments may have different available networks. By detecting, it can choose a network with good signal strength and high stability for connection, avoiding connection failure or unstable communication due to poor network signal.

[0013] More specifically, according to the detected network environment, the appropriate network connection method is selected. If there is a stable Wi-Fi hotspot around and the smart furniture supports Wi-Fi connection, it can be connected to the Wi-Fi network. If it is in a ZigBee network environment that has been built, the ZigBee module of the smart furniture can join the ZigBee network. Different network connection methods have different characteristics and application scenarios. Wi-Fi network has the characteristics of wide coverage and fast transmission speed, which is suitable for smart furniture with high data transmission requirements. ZigBee network has the characteristics of low power consumption and strong self-organizing network capability, which is suitable for the networking needs of some low-power devices. Selecting the appropriate network connection method can improve the performance and stability of the smart home network.

[0014] More specifically, if the selected network requires authentication information, such as the password of a Wi-Fi network, the user needs to input the corresponding authentication information through the interactive interface of the smart furniture (such as touch screen, keypad, etc.) or through a mobile phone APP, etc. After the network module receives the authentication information, it will try to connect to the network for verification. In order to ensure the security of the network, most networks have an authentication mechanism. Inputting the correct authentication information is a necessary condition for connecting to the network. Only by passing the authentication can a stable network connection be established.

[0015] More specifically, the network module uses the input authentication information to attempt to connect to the selected network. During the connection process, the network module communicates with the router or gateway in the network to exchange handshake information, assign IP addresses, and perform other operations. If the connection is successful, the network module will display the connection status as connected and can begin communicating with other smart home devices. Establishing a network connection is the core step in building a smart home network. Only when all smart home devices are successfully connected to the same local area network can they transmit and interact with each other, thereby realizing the various functions of the smart home.

[0016] More specifically, after a successful connection, each smart home device will send test data packets to each other to verify the stability and communication quality of the network connection. If a network connection problem is found, such as excessively high packet loss rate or excessive latency, the network module will automatically adjust some parameters, such as signal strength and communication frequency, or try to reconnect to other available networks. The stability and communication quality of the network connection directly affect the performance of the smart home system. By verifying and optimizing the network connection, potential problems can be identified and resolved in a timely manner, ensuring that the smart home network can operate stably and efficiently.

[0017] The second step involves collecting sound signals based on the smart home network and analyzing them to obtain the location and state characteristics of the source of the human voice.

[0018] Specifically, in the second step of the embodiment provided by this invention, each smart furniture unit within the smart home network continuously senses and records sound signals from the external environment to obtain the corresponding sound signals of each smart furniture unit. The microphones and other sound sensors built into each smart furniture unit continuously collect sound from the surrounding environment and convert it into electrical or digital signals for recording and storage. The sound signals are the basic data for subsequent analysis of the location and state of the source of human voices. The sound signals collected by smart furniture units in different locations will vary due to factors such as distance from the sound source and relative position. By collecting sound signals simultaneously from multiple smart furniture units, more comprehensive sound information can be obtained, providing rich data support for subsequent accurate analysis of the source of human voices.

[0019] More specifically, the sound signals collected by each piece of smart furniture are processed to separate human voices from ambient sounds in order to obtain the human voice signals within each sound signal. Signal processing techniques such as filtering and spectrum analysis can be used to separate human voices from background noise and other ambient sounds. Based on the pre-registered voice source account information, the voice signal is identified by its sound attribution characteristics to obtain the corresponding voice source information. For example, by comparing the timbre, pitch and other characteristics of the human voice, it can be determined which member of the family the voice belongs to.

[0020] More specifically, the sound performance characteristics of the human voice signals collected by each smart furniture are analyzed according to the human voice source information to obtain the sound performance characteristics of the human voice signals collected by each smart furniture corresponding to each human voice source. The sound performance characteristics include the volume characteristics and audio characteristics of the human voice signals corresponding to the specified human voice source. For example, the volume and frequency distribution of the human voice of the same person collected at different smart furniture are analyzed. Based on the setting position relationship between each smart furniture in the smart home network, the sound source positioning analysis is performed on the sound performance characteristics of the human voice signals collected by each smart furniture corresponding to each human voice source to obtain the position characteristics of each human voice source. For example, the position of the sound source is calculated according to the volume difference and time difference of the human voice of the same person collected at different smart furniture by using the triangular positioning principle.

[0021] More specifically, the environment sound will interfere with the analysis of human voice, and separating it can improve the accuracy of subsequent human voice feature analysis and more accurately focus on the sound emitted by the user. In a family environment, there may be multiple members, and identifying the ownership of the sound can clearly identify the identity of the sounder. This is very important for some personalized voice control services and more accurate analysis of the intentions of the sounder. The sound performance characteristics collected at different positions are different, and these characteristics contain information about the position of the sound source. By analyzing the volume, audio, and other characteristics, key data can be provided for subsequent positioning. Clear position characteristics of the human voice source help the smart home system better understand the user's use scenarios and needs. For example, if it is known that the user is speaking in a corner of the living room, the system can more specifically control the nearby smart devices.

[0022] More specifically, the sound signals collected by each smart furniture are recorded to generate the sound perception sequence of each smart furniture. The sound signals collected by each smart furniture within a period of time are arranged in chronological order to form a sequence. Based on the current time, the sound perception sequence of each smart furniture is analyzed for the sound frequency of each human voice source within a specified time range. At the same time, the fuzzy semantic recognition of each sound perception sequence is performed by using a pre-trained lightweight voice recognition model. For example, the number of times a member speaks within the past 5 minutes is counted, and the content of the sound is preliminarily understood semantically.

[0023] More specifically, the sound emission purpose of the human voice source is predicted according to the sound frequency analysis result and the fuzzy semantic recognition result to generate the corresponding state characteristics of each human voice source. For example, if the sound frequency of a member is high and the semantic content involves a request to switch the TV program, it can be predicted that the purpose of the sound emission is to control the TV, and the corresponding state characteristics are generated.

[0024] More specifically, sound perception sequences can reflect changes in sound over time, providing a complete data structure for subsequent frequency analysis and semantic recognition. This helps to more comprehensively analyze users' vocal behavior. Vocal frequency can indirectly reflect a user's activity state and the urgency of their needs, while fuzzy semantic recognition can provide a preliminary understanding of the semantic information in a user's voice. Combining the two allows for a more comprehensive grasp of the user's vocal intent, understanding the purpose and state characteristics of the voice's source. Smart home systems can then respond more intelligently; for example, if the system predicts that the user is resting, it can automatically adjust the brightness of lights and reduce the volume of devices.

[0025] The third step is to assign voice recognition tasks to the smart furniture specified in the smart home network based on the location features and the state features.

[0026] Specifically, in the third step of the embodiment provided by the present invention, the current voice scene is predicted based on the location and state features of each voice source to obtain voice scene prediction information. For example, if the location features show that the user is on the sofa in the living room, and the state features indicate that the user is in a relaxed state and the voice frequency is low, the current voice scene can be predicted as the user easily controlling the smart speaker in the living room to play music; if the location features show that the user is in the kitchen, and the state features show that the voice is rapid and the semantics involve cooking-related content, the prediction is that the user is issuing commands to control kitchen appliances during kitchen operations. Under different voice scenes, the user's voice needs and desired control objects are different. By predicting the voice scene, the user's intention can be understood more accurately, providing a basis for the subsequent reasonable allocation of voice recognition tasks and improving the accuracy and efficiency of voice control.

[0027] More specifically, based on the placement relationships of each smart furniture item in the smart home network and the device performance data of each smart furniture item, the sound scene prediction information is analyzed to determine the sound signal perception performance of each smart furniture item relative to the human voice source. This yields the perception performance parameters of each smart furniture item in the current sound scene. For example, factors such as the distance between the smart furniture and the sound source, whether there are obstacles in between, the microphone sensitivity of the device, and its anti-interference ability are considered. If a smart speaker is close to the sound source and has high microphone sensitivity, its perception performance parameters are higher. Different smart furniture items have different sound signal perception capabilities in different sound scenes. Analyzing the perception performance parameters can determine which smart furniture items are more suitable for receiving and recognizing the user's voice signal, thereby avoiding assigning the voice recognition task to devices that do not have good perception conditions and improving the success rate of voice recognition.

[0028] More specifically, by comparing the perception performance parameters of each smart furniture piece, corresponding voice recognition tasks can be assigned to each piece. Smart furniture pieces with high perception performance parameters can be assigned the main voice recognition tasks, while devices with slightly lower perception performance can serve as auxiliary devices for partial verification or backup recognition. For example, the smart speaker closest to the sound source and with good performance can be used as the main recognition device, responsible for complete voice content recognition; while a nearby smart TV can serve as an auxiliary device to perform simple verification of the recognition results. Reasonably allocating voice recognition tasks can make full use of the advantages of each smart furniture piece, improve the voice processing efficiency and accuracy of the entire smart home system, avoid all devices performing repetitive recognition work, reduce the waste of system resources, and reduce the risk of control command errors due to recognition failure of a single device.

[0029] The fourth step involves performing feature recognition on the collected sound signals based on the speech recognition task, generating voice control commands, and sending them to the corresponding smart furniture through the smart home network.

[0030] Specifically, in the fourth step of the embodiment provided by this invention, the sound signal acquisition mode and sound signal recognition mode of the specified smart furniture are adjusted based on the speech recognition task, so that the specified smart furniture can capture sound signals in the adjusted sound signal acquisition mode. For example, if the speech recognition task requires the focus on recognizing sounds in a specific frequency range, then the sound signal acquisition mode of the smart furniture will be adjusted to focus on collecting sounds in that frequency range. In terms of recognition mode, a recognition algorithm more suitable for the scenario is enabled, such as a noise reduction recognition algorithm for noisy environments. Different speech recognition tasks have different requirements for the acquisition and recognition of sound signals. By adjusting the acquisition and recognition modes, the smart furniture can more accurately capture and recognize task-related sound signals, improving the accuracy and efficiency of recognition. For example, in a noisy kitchen environment, adjusting the acquisition mode can enhance the capture of human voices, and adjusting the recognition mode can better filter out environmental noise.

[0031] More specifically, the system analyzes the sound signals using a modified sound signal recognition mode on each designated smart furniture piece to obtain the human voice interpretation information from each piece of smart furniture. The smart furniture then uses its built-in speech recognition algorithm to convert the collected sound signals into text information and performs preliminary semantic analysis on the text information to determine the approximate content expressed by the sound. The collected sound signals themselves are analog or digital audio data, which cannot be directly understood and processed by the smart home system. Only by analyzing the content of the sound signals and converting them into human voice interpretation information (text and semantic information) can a foundation be provided for subsequent judgment of smart home control needs.

[0032] More specifically, based on the voice recognition tasks assigned to each piece of smart furniture, corresponding voice interpretation weights are generated for each piece. The voice interpretation information of each piece of smart furniture is then weighted and fused based on these weights to generate comprehensive interpretation information. For example, the interpretation information of the smart furniture primarily responsible for voice recognition has a higher weight, while the interpretation information of the smart furniture assisting in recognition has a lower weight. The comprehensive interpretation information collected at each moment is arranged chronologically, and a correlation analysis is performed on the comprehensive interpretation information at each moment in the chronological arrangement. Based on the correlation analysis results, the comprehensive interpretation information at the current moment is verified. For example, if the user said "turn on the living room light" at the previous moment and said "dimming it a little" at the current moment, the correlation analysis can more accurately understand that the current instruction is directed at the living room light.

[0033] More specifically, when the information verification results show that the comprehensive interpretation information conforms to the smart furniture voice control standard, the control object and control form are analyzed based on the comprehensive interpretation information to serve as the smart furniture control requirements. For example, the comprehensive interpretation information may indicate that the control object is the living room air conditioner, and the control form is to raise the temperature by 2 degrees Celsius. Different smart furniture has different voice recognition capabilities and reliability. Through weighted fusion, the interpretation results of various smart furniture can be combined to improve the accuracy and reliability of the interpretation information. User voice commands often have coherence and contextual relevance. Through temporal arrangement and correlation analysis, the complete intent of the user can be better understood, avoiding misunderstandings caused by understanding a single command in isolation. Only by obtaining accurate control objects and control forms through verification and analysis can the generated voice control commands meet the user's real needs, enabling the smart home system to respond correctly.

[0034] More specifically, voice control commands are generated based on the smart home control requirements, and these commands are sent to the corresponding smart furniture via the smart home network. This allows the smart furniture to operate according to the voice control commands. For example, if the control requirement is to open the bedroom curtains, the system will generate corresponding control commands (such as specific codes or signals) and send them to the bedroom curtain smart controller via the network. Generating and sending control commands is the ultimate goal of achieving smart home voice control. Only by sending accurate control commands to the corresponding smart furniture can the smart furniture operate according to the user's wishes, thus realizing the automation and intelligent control of the smart home.

[0035] As can be seen from the above technical solution, the voice control method for smart homes provided by the present invention has the following beneficial effects: This invention constructs a network to achieve device interconnection and interoperability, analyzes the location and state characteristics of human voices, can accurately locate the speaker, grasp their intentions, rationally allocate voice recognition tasks, leverage the advantages of each smart furniture, improve recognition efficiency and accuracy, generate accurate control commands and transmit them stably, and realize effective control of smart furniture, allowing users to conveniently control their home with voice, improve the comfort and intelligence of life, and meet the needs of modern fast-paced life.

[0036] Furthermore, the steps of collecting sound signals based on the smart home network and analyzing them to obtain the location and state characteristics of the source of human voices include: Each smart home device within the smart home network continuously senses and records sound signals from the external environment to obtain the sound signals of each smart home device. Based on the positional relationship between the smart furniture pieces in the smart home network, the sound signals collected by each smart furniture piece are analyzed to locate the source of human voices, so as to obtain the location characteristics of the source of human voices. The purpose of sound emission is analyzed by collecting sound signals from each piece of smart furniture within a specified time range to obtain the state characteristics of the source of human voice.

[0037] Specifically, the sound sensors (such as microphones) in each smart furniture unit will be continuously working, sensing the sounds in the surrounding external environment in real time, converting the sensed sound signals into electrical or digital signals, and recording them according to a certain sampling frequency and precision to form sound signal data for each smart furniture unit. This data will be temporarily stored in the local storage device of the smart furniture for subsequent analysis. The sound signals are the basic data for subsequent analysis of the location and status characteristics of human voice sources. Only by comprehensively and accurately sensing and recording sound can useful information be extracted from this data. The sound signals collected by smart furniture in different locations will vary due to factors such as distance and angle from the sound source. Simultaneous collection by multiple smart furniture units can obtain richer sound information, providing more comprehensive data support for subsequent positioning and status analysis.

[0038] More specifically, the sound signals collected by each piece of smart furniture are filtered to remove noise interference, improve signal quality, and separate human voices from ambient sounds. For example, methods such as spectrum analysis and adaptive filtering are used to extract human voice signals from mixed sound signals. Based on the pre-registered voice account information of the human voice source, the separated human voice signals are identified by their sound attribution characteristics to determine which or which known speakers the human voice signals come from. The sound performance characteristics of each human voice source corresponding to the human voice signals collected by each piece of smart furniture are analyzed, such as volume characteristics (loudness of sound) and audio characteristics (frequency distribution of sound).

[0039] More specifically, based on the positional relationships between various smart furniture pieces in the smart home network (which can be determined through pre-configured coordinate information), and combined with the sound performance characteristics of each voice source corresponding to the voice signals collected by each smart furniture piece, algorithms such as triangulation and time difference of arrival (TDOA) are used to calculate the location characteristics of each voice source, i.e., the specific location of the speaker in space. Environmental noise can interfere with voice analysis; filtering and separation operations can improve the purity of the voice signal, making subsequent analysis more accurate. Identifying the sound attribution characteristics can distinguish different speakers, facilitating individual analysis of each speaker. The sound performance characteristics contain information related to the speaker's location, such as the volume being related to the distance from the sound source, and audio characteristics may also change due to different propagation paths. Clearly defining the location characteristics of the voice source helps the smart home system better understand the user's usage scenarios and needs. For example, if it is known that the user is making a sound in a corner of the living room, the system can more effectively control nearby smart devices and provide more personalized services.

[0040] More specifically, the sound signal recording and sequence generation involves: recording the sound signals collected from each piece of smart furniture in detail, generating a sound perception sequence for each piece of smart furniture in chronological order, and using the current time as a reference to analyze the frequency of each voice source within a specified time range (e.g., the past 5 minutes) of the sound perception sequence of each smart furniture, counting the number of voices and the time interval, etc., and using a pre-trained lightweight speech recognition model to perform fuzzy semantic recognition on each sound perception sequence to extract keywords and general semantic content from the sound. Based on the voice frequency analysis results and fuzzy semantic recognition results, combined with preset rules or machine learning models, the purpose of the voice source is predicted. For example, if the voice frequency is high and the semantics involve a request to switch TV programs, it can be predicted that the purpose of the voice is to control the TV, and then a corresponding state feature is generated for the voicer, such as "currently performing TV control operation".

[0041] More specifically, sound perception sequences can reflect changes in sound over time, providing a complete data structure for subsequent vocal frequency analysis and semantic recognition. This helps to analyze users' vocal behavior more comprehensively. Vocal frequency can reflect the user's activity state and the urgency of their needs, while fuzzy semantic recognition can initially understand the semantic information in the user's voice. The combination of the two can more comprehensively grasp the user's vocal intentions, understand the purpose and state characteristics of the source of the voice, and enable smart home systems to respond more intelligently. For example, if it is predicted that the user is in a resting state, the system can automatically adjust the brightness of the lights, reduce the volume of the devices, and improve the user experience.

[0042] Furthermore, based on the positional relationships between the smart furniture pieces in the smart home network, the step of performing location analysis on the sound signals collected by each smart furniture piece to obtain the location characteristics of the sound source includes: The sound signals collected by each piece of smart furniture are processed to separate human voice from ambient sound, so as to obtain the human voice signal in each sound signal; The voice signal is identified by its voice attribution feature based on the pre-registered voice account information to obtain the voice source information corresponding to the voice signal; wherein, the voice source information is used to describe the number of sources and the identity of the sources of the voice signal. Based on the information about the source of the human voice, the sound performance characteristics of the human voice signals collected by each piece of smart furniture are analyzed to obtain the sound performance characteristics of the human voice signals collected by each piece of smart furniture corresponding to each human voice source; wherein, the sound performance characteristics include the volume characteristics and audio characteristics of the human voice signals corresponding to the specified human voice source; Based on the positional relationship between the smart furniture pieces in the smart home network, the sound source localization analysis is performed on the human voice signals collected by each smart furniture piece, corresponding to the sound performance characteristics of each human voice source, to obtain the positional characteristics of each human voice source.

[0043] Specifically, a bandpass filter is used to preliminarily process the acquired sound signal, limiting the frequency range to the frequency range where human voices typically reside (generally 300Hz - 3400Hz) to remove high-frequency and low-frequency noise interference. An adaptive filtering algorithm is used to adjust the filter parameters according to the real-time characteristics of the sound signal, further reducing the impact of environmental noise. The ICA algorithm is used to separate the mixed sound signal into different independent components. By analyzing the characteristics of each component, the human voice component is identified and extracted, obtaining the human voice signal within each sound signal. Ambient noise can severely interfere with the analysis and localization of human voices. For example, the hum of an air conditioner or the noise of cars outside the window can blur the characteristics of the human voice signal, making it difficult to accurately identify and analyze. Through separation processing, the purity of the human voice signal can be improved, making subsequent sound attribution feature identification and localization analysis more accurate.

[0044] More specifically, feature extraction is performed on the separated voice signals. Commonly used features include fundamental frequency, formant frequency, and spectral envelope. These features reflect the unique timbre and pronunciation characteristics of the human voice. The extracted features are then matched with feature templates in the pre-registered voice source account information. Dynamic Time Warping (DTW) algorithms or deep learning-based feature matching methods can be used to calculate the similarity between the two. Based on the matching results, the source identity of the voice signal is determined. If the similarity exceeds a set threshold, the voice signal is considered to come from the corresponding registered person. At the same time, the number of different source identities is counted to obtain the number of voice signal sources and source identity information. In a smart home environment, there may be multiple different voice users. Identifying voice attribution features can distinguish different voice users, making it easier to analyze and locate each voice user individually. For example, different family members have different voice control habits and needs, and accurate identification can provide more personalized services.

[0045] More specifically, calculating the volume of the human voice signals collected by each piece of smart furniture corresponding to each human voice source can be achieved by statistically analyzing the amplitude of the human voice signals to obtain parameters such as average volume and maximum volume. Spectral analysis of the human voice signals can be performed to obtain their frequency distribution characteristics, such as calculating the energy distribution of different frequency bands, the location of formant frequencies, and bandwidth. These audio characteristics can reflect the pitch and timbre of the human voice. The sound performance characteristics are closely related to the location of the speaker. The volume and audio characteristics of the same speaker's voice collected by smart furniture in different locations will be different. For example, the closer the smart furniture is to the speaker, the louder the sound volume and the richer the high-frequency components of the sound. By analyzing these characteristics, important basis can be provided for subsequent sound source localization.

[0046] More specifically, based on the positional relationships between various smart furniture pieces in the smart home network, a spatial location model is established. The location of each smart furniture piece can be represented using Cartesian or polar coordinate systems. The sound characteristics of each sound source, captured by each smart furniture piece, are matched with the location model. For example, using the principle of triangulation, the location coordinates of the speaker are calculated based on the volume differences or arrival time differences of the same speaker captured by different smart furniture pieces. Through multiple calculations and verifications, combined with other auxiliary information (such as sound reflection and refraction), the final location characteristics of each sound source are determined. Clearly defining the location characteristics of the sound source is key to achieving precise control of the smart home. For instance, if it is known that a user is making a sound from the sofa in the living room, the smart home system can automatically adjust the brightness and angle of the living room lights, turn on nearby smart speakers to play music, etc. Through sound source localization analysis, the smart home system can better understand the user's usage scenarios and needs, providing more intelligent and personalized services.

[0047] Furthermore, the steps of analyzing the sound signals collected from each piece of smart furniture within a specified time range to obtain the state characteristics of the human voice source include: The sound signals collected by each piece of smart furniture are recorded to generate a sound perception sequence for each piece of smart furniture; Based on the current time, the sound perception sequence of each smart furniture piece is analyzed for the frequency of each human voice source within a specified time range. At the same time, a pre-trained lightweight speech recognition model is used to perform fuzzy semantic recognition on each of the sound perception sequences. Based on the results of frequency analysis and fuzzy semantic recognition, the purpose of human voice sources is predicted, and corresponding state features are generated for each human voice source.

[0048] Specifically, each smart furniture unit continuously collects sound signals and records them in chronological order. The recorded information includes basic characteristics such as the intensity and frequency of the sound signals, as well as the timestamp of the collection. These recorded data are organized into an ordered sequence, forming the sound perception sequence of each smart furniture unit. For example, the sound feature data collected within a fixed time interval (such as per second) are arranged sequentially. The sound perception sequence can completely present the changes in sound over time. It provides the basic data structure for subsequent sound frequency analysis and semantic recognition, enabling the analysis of sound from a time dimension and helping to discover patterns and regularities in the sound signals.

[0049] More specifically, using the current time as a baseline, a specified time range (e.g., the past 5 minutes) is defined. The sound perception sequences of each smart furniture piece are analyzed, and the number of times each voice source speaks within that time range is counted to calculate the frequency. For example, if a specific voice source speaks 10 times in these 5 minutes, its frequency is 2 times / minute. A pre-trained lightweight speech recognition model is used to process each sound perception sequence. This model can convert sound signals into text information and perform preliminary semantic understanding. Due to the diversity and uncertainty of speech, fuzzy semantic recognition is used here, meaning that precise semantic understanding is not pursued, but rather key semantic information and keywords are extracted. The frequency of speech can reflect the activity and urgency of the voice source. A higher frequency of speech indicates that the user is in a more urgent state or is engaged in frequent interaction, while a lower frequency of speech indicates that the user is in a more relaxed or quiet state. By recognizing the semantic information in the sound, the general intention of the user's speech can be understood. Although it is a fuzzy recognition, it can capture key information and provide important clues for subsequent prediction of the purpose of the speech. Moreover, lightweight speech recognition models can reduce the consumption of computing resources while ensuring a certain level of recognition accuracy, making them suitable for running on smart home devices.

[0050] More specifically, a series of preset rules are established to match the results of voice frequency analysis and fuzzy semantic recognition. The matching and recognition results are used to determine the purpose of the voice source in the current time period. The purpose of the voice includes communication between people, the sound of a movie or TV program, or the user's voice control of smart furniture. By judging the state characteristics of the current time period, the probability of the user making a voice control of smart furniture at the current time is analyzed, and corresponding state characteristics are generated. Based on these state characteristics, the probability of the currently acquired sound signal being a voice control can be determined, thereby adjusting the voice recognition level of each smart piece of furniture to cope with different working scenarios and avoid voice misjudgment.

[0051] Furthermore, the step of assigning voice recognition tasks to designated smart furniture in the smart home network based on the location features and the state features includes: Based on the location and state characteristics of each voice source, the current voice scene is predicted to obtain voice scene prediction information; Based on the placement relationship of each smart furniture in the smart home network and the device performance data of each smart furniture, the sound scene prediction information is analyzed to assess the sound signal perception performance of each smart furniture relative to the source of human voice, thereby obtaining the perception performance parameters of each smart furniture in the current sound scene. The perception performance parameters of each smart furniture piece are compared to assign corresponding voice recognition tasks to each piece of smart furniture.

[0052] Specifically, the system collects location and state characteristics of each voice source. Location characteristics pinpoint the speaker's exact location in space, while state characteristics reflect the speaker's purpose and emotional state. For example, location characteristics might indicate the user is on a sofa in the living room, while state characteristics suggest the user is relaxed and the speech involves music playback. This information is integrated to establish a predefined rule base for speech scenarios. The rule base contains speech scenarios corresponding to different combinations of location and state characteristics. The integrated information is matched against the rule base to find the speech scenario that best fits the current situation. For example, if the rule base defines "the user mentions music playback while relaxing on the sofa in the living room" as corresponding to "relaxed music playback scenario in the living room," then the current speech scenario can be predicted to be this scenario. Machine learning algorithms, such as decision trees and neural networks, are used to train the system on a large amount of historical data. This historical data contains different location and state characteristics and their corresponding actual speech scenarios. The current location and state characteristics are then input into the trained model, and the model outputs the predicted speech scenario.

[0053] More specifically, users have different voice needs and desired control objects in different speech scenarios. Accurately predicting the speech scenario can provide a foundation for the subsequent rational allocation of speech recognition tasks. For example, in a bedroom sleep scenario, users want to control smart devices at a low volume, mainly controlling sleep-related devices such as adjusting the brightness of bedside lamps and controlling air purifiers; while in a living room entertainment scenario, they need to control devices such as TVs and stereos for entertainment activities. By predicting the speech scenario, the system can better understand the user's intentions and improve the accuracy and efficiency of voice control.

[0054] More specifically, the system obtains the placement relationships of each smart furniture item from the configuration information of the smart home system. These relationships can be represented by coordinates or relative positions. For example, it records information such as the smart speaker being located in the center of the living room and the television being located on the front wall of the living room. The system also collects device performance data for each smart furniture item, including microphone sensitivity, frequency response range, and anti-interference capabilities. This data can be obtained through the device's technical specifications or its self-testing function.

[0055] More specifically, by combining the predicted information of the sound emission scene, the sound signal perception performance of each smart furniture relative to the source of human voice is analyzed. Factors considered include the distance between the smart furniture and the sound source, whether there are obstacles in between, and the performance parameters of the devices. For example, according to acoustic principles, the closer the device is to the sound source, the less obstacles there are in between, and the higher the microphone sensitivity, the better its sound signal perception performance. Through calculation and evaluation, the perception performance parameters of each smart furniture in the current sound emission scene are obtained, such as the perception score expressed in numerical form.

[0056] More specifically, different smart furniture pieces have varying abilities to perceive sound signals in different sound-generating scenarios. Understanding the perception performance parameters of each smart piece of furniture can help determine which devices are better suited to receive and recognize users' voice signals. For example, in a large living room, a smart speaker closer to the user is more likely to capture the user's voice clearly than a smart camera further away. Avoiding assigning voice recognition tasks to devices lacking good perception capabilities can improve the success rate of voice recognition and reduce false and missed recognitions.

[0057] More specifically, the perception performance parameters of each smart furniture item are compared and ranked from highest to lowest. For example, the perception scores of each smart furniture item are compared, with the device with the highest score ranked first. Based on the comparison results and a preset task allocation strategy, a corresponding voice recognition task is assigned to each smart furniture item. Common strategies include assigning the main voice recognition task to the device with the highest perception performance, with other devices serving as auxiliary verification or backup recognition; or assigning different parts of the voice recognition task to different devices based on the complexity of the task and the processing power of the devices. For example, for simple voice commands, the smart speaker with the best perception performance is used for recognition; for complex voice interactions, in addition to the smart speaker performing the main recognition, a nearby smart TV also participates in some semantic understanding and verification work.

[0058] More specifically, the assigned voice recognition tasks are communicated to the corresponding smart home devices, and necessary configurations are performed. For example, task instructions are sent to the smart devices, informing them of the types of voices to be recognized, task priorities, and other information. This allows the devices to work according to the assigned tasks. Reasonable allocation of voice recognition tasks can fully utilize the advantages of each smart home device, improve the voice processing efficiency and accuracy of the entire smart home system, avoid repetitive recognition work by all devices, reduce waste of system resources, and lower the risk of control command errors due to recognition failures by a single device. By optimizing task allocation, the smart home system can respond to user voice commands more intelligently and efficiently.

[0059] Furthermore, the steps of performing feature recognition on the collected sound signals based on the speech recognition task, generating voice control commands, and sending them to the corresponding smart furniture through the smart home network include: Based on the speech recognition task, the sound signal acquisition mode and sound signal recognition mode of the specified smart furniture are adjusted so as to capture sound signals through each specified smart furniture in the adjusted sound signal acquisition mode. The sound signal is analyzed by using the adjusted sound signal recognition mode on each designated smart furniture to obtain the human voice interpretation information obtained from the analysis of each smart furniture. Based on the voice recognition tasks assigned to each piece of smart furniture, the decoded human voice information is comprehensively analyzed to determine the current smart home control needs of the sound signal; Voice control commands are generated based on the smart home control requirements, and the voice control commands are sent to the corresponding smart furniture through the smart home network so that the smart furniture can operate according to the voice control commands.

[0060] Specifically, carefully study the speech recognition tasks assigned to the designated smart furniture, clarifying the specific requirements of the task, such as the type of speech to be recognized (e.g., command speech, interrogative speech), the key frequency range of sound to focus on, and the required recognition accuracy. Based on the task requirements, adjust the parameters of the smart furniture's sound acquisition equipment (e.g., microphone). For example, if the task focuses on low-frequency sounds, adjust the microphone gain to enhance the ability to capture low-frequency signals; if the environment is noisy, enable noise reduction or adjust the sampling frequency to improve the quality of sound acquisition. Adjust the speech recognition algorithms and models built into the smart furniture. For example, if the task is to recognize the speech of a specific user, load the user's speech feature template to improve recognition accuracy; if the task requires a fast response, optimize the complexity of the recognition algorithm to reduce recognition time. Different speech recognition tasks have different requirements for sound signal acquisition and recognition. By adjusting the acquisition and recognition modes, the smart furniture can more accurately capture and recognize task-related sound signals, improving recognition accuracy and efficiency. For example, the requirements for sound acquisition and recognition differ in a quiet bedroom environment and a noisy living room environment; adjusting the mode allows the device to adapt to different scenarios.

[0061] More specifically, the collected sound signals undergo preprocessing, including noise removal, filtering, and normalization, to improve signal quality and facilitate subsequent analysis. Features, such as acoustic features (pitch, duration, timbre, etc.) and linguistic features (vocabulary, grammar, etc.), are extracted from the preprocessed sound signals. These features form the basis for speech recognition and semantic understanding. Using the adjusted sound signal recognition mode, the extracted features are matched with the speech recognition model to convert the sound signal into text information. Simultaneously, preliminary semantic analysis is performed on the text information to determine the approximate content expressed by the sound, obtaining the human voice interpretation information from each smart home device. The collected sound signals themselves are analog or digital audio data, which cannot be directly understood and processed by the smart home system. By analyzing the content of the sound signals, they are converted into human voice interpretation information (text and semantic information), providing a foundation for subsequent judgment of smart home control needs.

[0062] More specifically, based on the voice recognition tasks assigned to each smart furniture piece, corresponding voice interpretation weights are generated for each piece. For example, the interpretation information of the smart furniture primarily responsible for voice recognition has a higher weight, while the interpretation information of the smart furniture assisting in recognition has a lower weight. Based on the assigned weights, the voice interpretation information of each smart furniture piece is weighted and fused to generate comprehensive interpretation information. For example, the interpretation results of multiple smart furniture pieces on the same voice segment are weighted and averaged or voted to obtain a more accurate interpretation. The comprehensive interpretation information collected at each moment is arranged chronologically, and the comprehensive interpretation information at each moment in the chronological arrangement is analyzed for correlation. Combining historical voice commands and the current state of the smart home system, the rationality of the comprehensive interpretation information at the current moment is verified. For example, if the user said "turn on the living room light" at the previous moment, and said "dimming it" at the current moment... "Point" refers to the need to dim the living room lights, which can be determined through correlation analysis. When the information verification results show that the comprehensive interpretation information conforms to the smart furniture voice control standard, the control object (such as the living room air conditioner, bedroom curtains, etc.) and control form (such as opening, closing, adjusting temperature, etc.) are obtained based on the comprehensive interpretation information analysis. As a smart furniture control requirement, different smart furniture has different voice recognition capabilities and reliability. Through weighted fusion, the interpretation results of each smart furniture can be combined to improve the accuracy and reliability of the interpretation information. Voice commands often have coherence and contextual relevance. Through correlation analysis and verification, the complete intent of the user can be better understood, avoiding misunderstandings caused by understanding a single command in isolation. Accurately judging the smart furniture control needs is the key to achieving precise control of smart homes. Only by clarifying the control object and control form can the correct voice control command be generated.

[0063] More specifically, based on the control needs of smart homes, corresponding voice control commands are generated. The format and content of the commands must conform to the communication protocol of the smart home network and the control interface requirements of the smart furniture. For example, if the control requirement is to open the bedroom curtains, the generated command contains a data packet with specific codes and parameters. The generated voice control command is sent to the corresponding smart furniture through the smart home network. During transmission, the accuracy and integrity of the command must be ensured, and encryption and verification technologies can be used. After receiving the voice control command, the corresponding smart furniture parses and verifies the command, and then performs the operation according to the command's requirements. For example, after receiving the open command, the smart curtain controller drives the motor to open the curtains. Generating and sending control commands is the ultimate goal of realizing smart home voice control. Only by sending accurate control commands to the corresponding smart furniture can the smart furniture operate according to the user's wishes, realizing the automation and intelligent control of the smart home.

[0064] Furthermore, the steps for comprehensively analyzing the parsed human voice interpretation information based on the voice recognition tasks assigned to each smart furniture piece to determine the current smart home control needs of the sound signal include: Based on the voice recognition task assigned to each piece of smart furniture, a corresponding voice interpretation weight is generated for each piece of smart furniture. The voice interpretation information of each piece of smart furniture is then weighted and fused based on the voice interpretation weight to generate comprehensive interpretation information. The comprehensive interpretation information collected at each time point is arranged in chronological order, and the comprehensive interpretation information at each time point in the chronological order is analyzed for correlation before and after, so as to verify the comprehensive interpretation information at the current time point based on the results of the correlation analysis. When the information verification results show that the comprehensive interpretation information conforms to the smart furniture voice control standard, the control object and control form are obtained based on the comprehensive interpretation information, which serve as the smart furniture control requirements.

[0065] Specifically, the speech recognition tasks assigned to each piece of smart furniture are analyzed in depth, taking into account the importance and difficulty of the tasks and the role of each smart furniture in the entire speech recognition process. For example, smart furniture that undertakes the main recognition task has high task importance; if the task involves complex semantic understanding, the difficulty is greater. Based on the task evaluation results, corresponding human voice interpretation weights are generated for each piece of smart furniture. The weight can be a value between 0 and 1. The higher the weight, the greater the proportion of the human voice interpretation information of that smart furniture in the comprehensive analysis. For example, the weight of the smart speaker that is mainly responsible for recognition and has good performance can be set to 0.8, and the weight of the smart camera that assists in recognition can be set to 0.2.

[0066] More specifically, the voice interpretation information of each smart furniture piece is multiplied by its corresponding weight, and then the results are summed to obtain the comprehensive interpretation information. For example, if the smart speaker's interpretation information is "turn on the TV," with a weight of 0.8, and the smart camera's interpretation information is "turn on the TV," with a weight of 0.2, then the comprehensive interpretation information is the result of fusing these two interpretation information according to their weights. Different smart furniture pieces have varying capabilities and reliability in the voice recognition process. By generating voice interpretation weights and performing weighted fusion, the characteristics and functions of each smart furniture piece can be fully considered, improving the accuracy and reliability of the comprehensive interpretation information. This avoids deviations in overall judgment due to misidentification of individual smart furniture pieces, making the comprehensive interpretation information more reflective of the user's true intent.

[0067] More specifically, the comprehensive interpretation information collected at each moment is arranged chronologically to form a time series. This clearly shows how the comprehensive interpretation information changes over time. Analyzing the comprehensive interpretation information at each moment in the time series allows us to examine the correlation between interpretation information at adjacent moments or within a certain time range. For example, if a user said "turn on the living room light" at one moment and said "dim it a little" at the current moment, correlation analysis can determine that the current command is to dim the living room light. Based on the results of the correlation analysis, the comprehensive interpretation information at the current moment is verified to check whether the information is logical and contextual, and whether it matches the current state of the smart home system. For example, if the living room light is currently off, but the comprehensive interpretation information is "dim the living room light," then the information is illogical and needs further verification. User voice commands often have coherence and contextual relevance. Through chronological arrangement and correlation analysis, we can better understand the user's complete intent and avoid misunderstandings caused by interpreting a single command in isolation. Information verification ensures the rationality and accuracy of the comprehensive interpretation information, improving the reliability of judging smart home control needs.

[0068] More specifically, when the information verification results show that the comprehensive interpretation information conforms to the smart furniture voice control standard, the comprehensive interpretation information is matched with pre-set rules. These rules define the control object and control form corresponding to different semantic information. For example, the rule stipulates that the control object corresponding to "turn on + TV" is the TV, and the control form is "turn on". For some complex comprehensive interpretation information, it may not be possible to directly determine the control object and control form through rule matching. Semantic understanding and reasoning are required. Natural language processing technology is used to analyze the vocabulary, grammar and semantic relationships in the comprehensive interpretation information to infer the user's true intention. For example, for "turn on the device in the bedroom that can adjust the temperature", semantic understanding can infer that the control object is the bedroom air conditioner and the control form is "turn on".

[0069] More specifically, based on the results of rule matching and semantic understanding reasoning, the control object and control form are determined and used as the control requirements for smart furniture. For example, the control object is determined to be the living room stereo, and the control form is to play music. Clarifying the control object and control form is the key to realizing voice control of smart homes. Only by accurately analyzing the user's control needs can the correct voice control commands be generated and sent to the corresponding smart furniture, so that the smart home system can operate according to the user's wishes and achieve the purpose of intelligent control.

[0070] Based on the technical content of the smart home voice control method described in the above-disclosed embodiments, the present invention provides a smart home voice control system, the structure of which is as follows: Figure 2 The voice control method for implementing the smart home as described in any one of the first aspects includes: The network connection module is used to connect each smart furniture piece to a local area network through the network module set in each smart furniture piece, so as to build a smart home network; The feature analysis module is used to collect sound signals based on the smart home network and analyze them to obtain the location and state features of the source of the human voice; The task allocation module is used to allocate voice recognition tasks to specified smart furniture in the smart home network based on the location features and the status features. The voice control module is used to perform feature recognition on the collected sound signals based on the voice recognition task, generate voice control commands, and send them to the corresponding smart furniture through the smart home network.

[0071] In this embodiment, the specific implementation of each module in the above system embodiment is described in the above method embodiment, and will not be repeated here.

[0072] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0073] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0074] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voice control method for smart homes, characterized in that, include: By connecting the smart furniture pieces to a local area network through network modules installed within each piece, a smart home network can be built. Based on the smart home network, the location and state characteristics of the source of human voices are obtained by collecting and analyzing sound signals. Based on the location features and the state features, assign voice recognition tasks to the smart furniture specified in the smart home network; Based on the voice recognition task, the collected sound signals are characterized, voice control commands are generated, and sent to the corresponding smart furniture through the smart home network.

2. The voice control method for smart homes as described in claim 1, characterized in that, The steps for collecting sound signals based on the smart home network and analyzing them to obtain the location and state characteristics of the human voice source include: Each smart home device within the smart home network continuously senses and records sound signals from the external environment to obtain the sound signals of each smart home device. Based on the positional relationship between the smart furniture pieces in the smart home network, the sound signals collected by each smart furniture piece are analyzed to locate the source of human voices, so as to obtain the location characteristics of the source of human voices. The purpose of sound emission is analyzed by collecting sound signals from each piece of smart furniture within a specified time range to obtain the state characteristics of the source of human voice.

3. The voice control method for smart homes as described in claim 2, characterized in that, Based on the positional relationships between smart furniture pieces in the smart home network, the steps for performing location analysis on the sound signals collected by each smart furniture piece to obtain the location characteristics of the sound source include: The sound signals collected by each piece of smart furniture are processed to separate human voice from ambient sound, so as to obtain the human voice signal in each sound signal; The voice signal is identified by its voice attribution feature based on the pre-registered voice account information to obtain the voice source information corresponding to the voice signal; wherein, the voice source information is used to describe the number of sources and the identity of the sources of the voice signal. Based on the information about the source of the human voice, the sound performance characteristics of the human voice signals collected by each piece of smart furniture are analyzed to obtain the sound performance characteristics of the human voice signals collected by each piece of smart furniture corresponding to each human voice source; wherein, the sound performance characteristics include the volume characteristics and audio characteristics of the human voice signals corresponding to the specified human voice source; Based on the positional relationship between the smart furniture pieces in the smart home network, the sound source localization analysis is performed on the human voice signals collected by each smart furniture piece, corresponding to the sound performance characteristics of each human voice source, to obtain the positional characteristics of each human voice source.

4. The voice control method for smart homes as described in claim 2, characterized in that, The steps for analyzing the sound signals collected from each piece of smart furniture within a specified time range to obtain the state characteristics of the human voice source include: The sound signals collected by each piece of smart furniture are recorded to generate a sound perception sequence for each piece of smart furniture; Based on the current time, the sound perception sequence of each smart furniture piece is analyzed for the frequency of each human voice source within a specified time range. At the same time, a pre-trained lightweight speech recognition model is used to perform fuzzy semantic recognition on each of the sound perception sequences. Based on the results of frequency analysis and fuzzy semantic recognition, the purpose of human voice sources is predicted, and corresponding state features are generated for each human voice source.

5. The voice control method for smart homes as described in claim 1, characterized in that, The steps of assigning speech recognition tasks to specified smart furniture in the smart home network based on the location features and the state features include: Based on the location and state characteristics of each voice source, the current voice scene is predicted to obtain voice scene prediction information; Based on the placement relationship of each smart furniture in the smart home network and the device performance data of each smart furniture, the sound scene prediction information is analyzed to assess the sound signal perception performance of each smart furniture relative to the source of human voice, thereby obtaining the perception performance parameters of each smart furniture in the current sound scene. The perception performance parameters of each smart furniture piece are compared to assign corresponding voice recognition tasks to each piece of smart furniture.

6. The voice control method for smart homes as described in claim 5, characterized in that, The steps of performing feature recognition on the collected sound signals based on the speech recognition task, generating voice control commands, and sending them to the corresponding smart furniture through the smart home network include: Based on the speech recognition task, the sound signal acquisition mode and sound signal recognition mode of the specified smart furniture are adjusted so that subsequent sound signals can be captured through each specified smart furniture in the adjusted sound signal acquisition mode. The sound signals are analyzed using the adjusted sound signal recognition mode for each designated smart furniture piece to obtain the human voice interpretation information corresponding to each smart furniture piece. Based on the voice recognition tasks assigned to each piece of smart furniture, the decoded human voice interpretation information is comprehensively analyzed to determine the current smart home control needs of the sound signal; Voice control commands are generated based on the smart home control requirements, and the voice control commands are sent to the corresponding smart furniture through the smart home network so that the smart furniture can operate according to the voice control commands.

7. The voice control method for smart homes as described in claim 6, characterized in that, The steps for comprehensively analyzing the decoded human voice information based on the voice recognition tasks assigned to each smart furniture piece to determine the current smart home control needs of the sound signal include: Based on the voice recognition task assigned to each piece of smart furniture, a corresponding voice interpretation weight is generated for each piece of smart furniture. The voice interpretation information of each piece of smart furniture is then weighted and fused based on the voice interpretation weight to generate comprehensive interpretation information. The comprehensive interpretation information collected at each time point is arranged in chronological order, and the comprehensive interpretation information at each time point in the chronological order is analyzed for correlation before and after, so as to verify the comprehensive interpretation information at the current time point based on the results of the correlation analysis. When the information verification results show that the comprehensive interpretation information conforms to the smart furniture voice control standard, the control object and control form are obtained based on the comprehensive interpretation information, which serve as the smart furniture control requirements.

8. A voice control system for smart homes, characterized in that, include: The network connection module is used to connect each smart furniture piece to a local area network through the network module set in each smart furniture piece, so as to build a smart home network; The feature analysis module is used to collect sound signals based on the smart home network and analyze them to obtain the location and state features of the source of the human voice; The task allocation module is used to allocate voice recognition tasks to specified smart furniture in the smart home network based on the location features and the status features. The voice control module is used to perform feature recognition on the collected sound signals based on the voice recognition task, generate voice control commands, and send them to the corresponding smart furniture through the smart home network.