Whole house intelligent scene generation system based on natural language

Through a whole-house intelligent scene generation system based on natural language, the user behavior and environmental status are captured in real time, combined with in-depth semantic analysis and dynamic feedback, the problem of insufficient personalized understanding of user needs in the existing technology is solved, and personalized scenario configuration and user experience improvement is achieved.

CN120010282AInactive Publication Date: 2025-05-16SHENZHEN UASCENT TECH CO LTD

Patent Information

Application Number
CN202510497076.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing whole-house intelligent scenario generation technology lacks personalized understanding of user needs, which leads to the inability to match personalized needs and the user experience is poor.

Method used

A whole-house intelligent scene generation system based on natural language is adopted to capture user behavior habits and environmental states in real time through scene perception components, combine the in-depth semantic analysis of the component with semantic understanding components, generate personalized scene configuration solutions, and continuously optimize the scene generation strategy through dynamic feedback components.

Benefits of technology

It achieves accurate understanding and personalized response to user needs, improves user experience, and enhances the adaptability and intelligence level of smart home systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010282A_ABST
    Figure CN120010282A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, particularly provides a whole-house intelligent scene generation system based on a natural language, and solves the problem that the experience feeling of a user on whole-house intelligence needs to be further improved. According to the system, a scene sensing assembly captures behavior habits and environment states of a user in real time through multi-mode data of a voice sensor, an image sensor, an environment sensor and the like; the semantic understanding component performs deep semantic analysis on the user voice instruction by using a large language model, obtains associated behavioral habits and ring states according to a deep semantic analysis result, and generates a scene configuration scheme in combination with the real-time state of the whole house equipment; and the dynamic feedback component dynamically optimizes a scene generation strategy by continuously collecting user feedback and environment data, and continuously adjusts scene configuration according to the satisfaction degree of the user, the equipment operation state and the external environment change. According to the embodiment of the invention, a complete closed loop from perception, understanding to optimization is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a whole-house intelligent scene generation system based on natural language. Background Art

[0002] In the field of whole-house smart scene generation, the background technology mainly involves the integration and automatic control technology of smart home systems. With the rapid development of Internet of Things (IoT) technology, smart home systems are gradually changing from the control of a single device to the collaborative work of multiple devices. Whole-house smart scene generation technology aims to achieve user-defined scene modes by integrating multiple smart devices (such as smart lighting, temperature control systems, security equipment, audio systems, etc.), thereby improving the quality of life and energy efficiency.

[0003] However, the existing technology lacks personalized understanding of user needs, resulting in the inability to match smart scenes with personalized needs, making the user experience of whole-house intelligence need to be further improved. Summary of the invention

[0004] In order to achieve the above object, the present invention adopts the following technical scheme: In one aspect of the present invention, a whole-house intelligent scene generation system based on natural language is provided, comprising: The scene perception component is responsible for capturing the user's behavior habits and environmental status in real time through multimodal data from voice sensors, image sensors, and environmental sensors; and associating the user's instructions collected by the voice sensor with the behavior habits and environmental status; The semantic understanding component is responsible for using a large language model to perform deep semantic analysis on user voice commands. Based on the results of the deep semantic analysis, it obtains the associated behavioral habits and environmental status, and generates a scenario configuration plan based on the real-time status of all the equipment in the house. The dynamic feedback component is responsible for dynamically optimizing the scene generation strategy by continuously collecting user feedback and environmental data, and continuously adjusting the scene configuration based on user satisfaction, device operating status, and external environment changes.

[0005] In an optional implementation, the scene perception component includes: The speech decomposition module is responsible for extracting features from speech commands and decomposing them into feature vectors in multiple dimensions, such as time, space, and semantics. At the same time, the environmental state data collected by the environmental sensor is mapped into an environmental state matrix in real time. The feature vector of the speech command is dynamically matched with the environmental state matrix through adaptive alignment to establish a preliminary association relationship. The graph construction module is responsible for constructing a behavior feature graph based on user behavior habit data, including multi-dimensional information such as time series, spatial distribution, and intensity change; through behavior-command association, the feature vector of the voice command is deeply coupled with the behavior feature graph; The weight calculation module is responsible for real-time analysis of the correlation between the current environment status and historical behavior patterns. It assigns different weights to each feature dimension of the voice command through dynamic weight allocation. The weights are dynamically adjusted according to changes in the environment status and the evolution of behavioral habits. The vector generation module is responsible for multimodal data fusion of the feature vector of voice commands, the environmental state matrix and the behavioral feature map to generate an enhanced semantic vector that includes environmental adaptability and behavioral habits.

[0006] In an optional implementation, the speech decomposition module includes: The multi-dimensional analysis submodule is responsible for decomposing the voice command into feature vectors of time, space and semantic dimensions after feature extraction, which respectively represent the information of the voice command in different dimensions; the environmental status data is collected through environmental sensors and mapped into an environmental status matrix, which contains the real-time status information of the current environment; The hierarchical mapping submodule is responsible for dividing the environmental data into the basic layer, dynamic layer and associated layer through hierarchical mapping; The dynamic matching submodule is responsible for preliminarily matching the temporal, spatial and semantic feature vectors of the voice command with the basic layer of the environment state matrix to find possible association points; through the analysis of the association layer, the semantic dimension of the voice command is deeply matched with the association information of the environment state matrix.

[0007] In an optional implementation, the time dimension of the multi-dimensional analysis submodule captures the time characteristics of the voice command; the spatial dimension reflects the spatial characteristics of the voice command; and the semantic dimension analyzes the semantic content of the voice command.

[0008] In an optional implementation, the base layer of the hierarchical mapping submodule contains basic physical parameters of the environment; the dynamic layer reflects real-time changes in environmental states; and the correlation layer captures correlations between environmental states.

[0009] In an optional implementation, the graph construction module comprises: The feature spatiotemporal decomposition submodule is responsible for decomposing the user behavior habit data into multiple time series segments according to the time dimension to capture the temporal characteristics of the user behavior; decomposing the user behavior habit data into multiple spatial distribution areas according to the spatial dimension to capture the spatial characteristics of the user behavior; decomposing the user behavior habit data into multiple intensity change curves according to the intensity dimension to capture the intensity characteristics of the user behavior; The feature hierarchical modeling submodule is responsible for building a basic feature model of user behavior based on time series, spatial distribution and intensity change data; capturing real-time changes and trends in user behavior to form a dynamic feature model; analyzing the correlation between user behaviors to form a correlation feature model. The deep coupling submodule is responsible for mapping the feature vectors of voice commands with the behavioral feature map to find the potential correlation between commands and behaviors. Through the hierarchical model of the behavioral feature map, it identifies the user behavior pattern under the current environmental state and deeply couples it with the voice commands.

[0010] In an optional implementation, the weight calculation module includes: The data mapping submodule is responsible for collecting the current environmental status data through environmental sensors and mapping it into an environmental status matrix. At the same time, the voice decomposition module extracts features from the user's voice commands and generates feature vectors containing multiple dimensions of time, space and semantics. The feature recognition submodule is responsible for analyzing the correlation between the current environment state and the user behavior pattern by combining the historical behavior feature map; by comparing the current environment state with the historical behavior pattern, it identifies which feature dimensions are more important in the current scenario; The weight assignment submodule is responsible for dynamically assigning weights to each feature dimension of the voice command.

[0011] In an optional implementation, the semantic understanding component includes: The model generation module is responsible for using a semantic analysis mechanism to convert the user's natural language instructions into executable scene configuration solutions; through deep semantic association, combined with the real-time status of all the equipment in the house, a dynamic scene model is generated; The model triggering module is responsible for receiving voice commands and triggering the large language model; The solution execution module is responsible for generating and executing highly personalized scenario solutions through perception and understanding mechanisms.

[0012] In an optional implementation, the model generation module includes: The multimodal data processing submodule is responsible for converting the user's natural language instructions into structured semantic information, extracting key actions, objects and context information in the instructions; capturing the user's behavior patterns in real time and converting them into time series features; capturing the current environment state in real time and quantifying it into a computable feature vector; combining the user's historical behavior data and environmental change trends to form a multi-dimensional data foundation; The semantic association and demand inference submodule is responsible for deeply associating voice commands with user behavior patterns to determine whether the behavior conforms to the typical pattern of commands; inferring changes in user needs by analyzing the association between user behavior patterns and environmental changes; and analyzing the association between voice commands and environmental changes to optimize the execution effect of commands; The dynamic scene model generation submodule is responsible for generating a state model of the current scene based on multimodal data, including the user's behavior status, device operation status, and environmental parameters; integrating historical behavior data and environmental change trends into the scene model to form a deep understanding of user habits and preferences; and predicting possible changes in user needs by analyzing the current scene status and historical data, and generating corresponding prediction information.

[0013] In an optional implementation, the dynamic scene model generation submodule includes: The association analysis unit is responsible for analyzing the association between user behavior patterns and environmental changes. The association analysis is based on current behavior and environmental data, combined with historical data to form a deep understanding of user habits. The data integration unit is responsible for generating a state model of the current scene based on the collected multimodal data, including the user's behavior status, device operation status, and environmental parameters; by integrating historical behavior data and environmental change trends into the model; The result prediction unit is responsible for predicting possible changes in user needs by analyzing the current scene status and historical data; the prediction is based on current behavior and environmental data, and the user's historical habits; it generates corresponding prediction information and adjusts environmental parameters in advance to adapt to changes in user needs.

[0014] The scene perception component of the present invention constructs a comprehensive user behavior and environmental portrait through voice, image, environmental sensors, etc.; continuously captures user behavior patterns and preferences, and establishes a dynamic user model; establishes intelligent associations between voice commands and user habits and environmental status; provides a personalized service foundation: provides accurate data support for scene configuration; makes smart homes more natural and more intimate. The semantic understanding component accurately understands the deep meaning of user commands; combines commands with user habits, environmental status, and device status; outputs the most optimized scene configuration suggestions; realizes natural interaction: allows users to control their homes in the most natural way; enables the system to actively provide the best solution; and enables the home system to truly have the ability to "think". The dynamic feedback component continuously optimizes the scene strategy through user feedback, dynamically adjusts according to environmental changes and user satisfaction, and realizes continuous improvement of scene configuration; realizes system evolution, allowing the smart home system to continue to grow; provides personalized services, allowing the system to understand users more and more; improves system reliability, and ensures that the scene configuration is always optimal. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a block diagram of a natural language-based whole-house smart scene generation system provided in Example 1 of the present invention; Figure 2 This is a block diagram of the scene perception component provided in Embodiment 2 of the present invention; Figure 3 This is a block diagram of a speech decomposition module provided in Embodiment 3 of the present invention; Figure 4 This is a block diagram of a graph construction module provided in Example 4 of the present invention; Figure 5 This is a block diagram of a weight calculation module provided in Embodiment 5 of the present invention; Figure 6 This is a block diagram of the semantic understanding component provided in Embodiment 6 of the present invention; Figure 7 This is a block diagram of a model generation module provided in Embodiment 7 of the present invention; Figure 8 A block diagram of a dynamic scene model generation submodule provided in Embodiment 8 of the present invention; Fig. 9 A block diagram of an electronic device provided by the present invention; Fig.10 A block diagram of a computer-readable storage medium provided for the present invention. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present invention will be described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0017] In the following, the terms "first", "second", etc. are used only for convenience of description and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0018] In the present invention, unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense, for example, "connection" can be a fixed mechanical connection, or a detachable mechanical connection, or integrated; or, "connection" can be a direct connection, or an indirect connection through an intermediate medium. In addition, unless otherwise clearly specified and limited, the term "coupling" should be understood in a broad sense, for example, "coupling" can be a direct electrical connection, such as physical contact and electrical conduction between two components, and can also be understood as electrical connection between different components in a circuit structure through physical lines such as copper foil or wires on a printed circuit board (PCB) that can transmit electrical signals to transmit electrical signals; or, "coupling" can be an indirect electrical connection between two components through an intermediate medium; or, "coupling" can be an electrical connection between two components in an air-spaced / non-contact manner, for example, two components are electrically connected by capacitive coupling to transmit electrical signals.

[0019] In the embodiments of the present invention, directional terms such as "up", "down", "left" and "right" may be defined including but not limited to the orientation relative to the schematic placement of the components in the drawings. It should be understood that these directional terms may be relative concepts, which are used for relative description and clarification, and may change accordingly according to the change of the orientation of the components in the drawings.

[0020] The embodiments of the present invention can be used in smart home scenarios and personalized life scenario generation. When a user issues a natural language command at home (such as "prepare dinner" or "I want to watch a movie"), the system captures the user's voice command, behavior pattern, and environmental changes through the scene perception component, and generates a personalized scene solution in combination with the deep semantic analysis of the semantic understanding component. For example: in the "prepare dinner" scenario, the system will automatically turn on the kitchen lighting, adjust the oven temperature, and recommend recipes based on the user's historical habits; in the "watch a movie" scenario, the system will automatically close the curtains, dim the lights, turn on the TV, and recommend movies based on the user's viewing preferences. Dynamic scene optimization, through the dynamic feedback component, real-time monitoring of user behavior feedback and environmental changes, dynamically optimize the scene configuration. For example: if the user gets up and leaves during the movie, the system will pause the playback and brighten the lights; if the user is not satisfied with the recommended movie, the system will adjust the recommendation strategy based on the feedback.

[0021] Smart office scene, efficient office scene generation. When the user issues a natural language command in the office (such as "start a meeting" or "I need to focus on work"), the system will generate the corresponding office scene according to the command. For example: in the "start a meeting" scene, the system will automatically turn on the projection equipment in the conference room, adjust the light brightness, and adjust the air conditioning temperature according to the number of participants; in the "focus on work" scene, the system will automatically turn off irrelevant equipment, dim the ambient light, and play white noise to improve concentration. Dynamic environment adjustment, through the dynamic feedback component, the user's office status and environmental changes are perceived in real time, and the office scene is dynamically adjusted. For example: if it is detected that the user has been inactive for a long time, the system will remind the user to take a break and adjust the ambient light; if the temperature in the conference room is too high, the system will automatically lower the air conditioning temperature.

[0022] Smart hotel scenarios, personalized check-in experience. When users issue natural language commands in the hotel room (such as "I want to rest" or "I need a wake-up service"), the system will generate a personalized check-in scenario based on the command. For example: in the "rest" scenario, the system will automatically close the curtains, dim the lights, adjust the air-conditioning temperature, and play sleep-inducing music; in the "wake-up service" scenario, the system will generate a personalized wake-up plan based on the user's historical preferences (such as liking a gentle wake-up). Dynamic service optimization, through the dynamic feedback component, collect user feedback and environmental data in real time to optimize the service experience. For example: if the user is not satisfied with the room temperature, the system will adjust the air-conditioning settings based on the feedback; if the user uses a service (such as ordering food) many times, the system will actively recommend related services.

[0023] Smart elderly care scenarios, health monitoring and scenario generation. When elderly users issue natural language commands (such as "I feel a little cold" or "I want to listen to the news"), the system will generate corresponding health monitoring scenarios based on the commands. For example: in the "feeling cold" scenario, the system will automatically increase the indoor temperature and remind the user to add clothes; in the "listening to the news" scenario, the system will play relevant news content based on the user's interest preferences; dynamic health management, through dynamic feedback components, real-time monitoring of the health status and environmental changes of elderly users, and dynamic adjustment of scenario configuration. For example: if it is detected that the user has not been active for a long time, the system will remind the user to get up and move around; if the indoor air quality is poor, the system will automatically turn on the air purifier.

[0024] Smart education scenarios, personalized learning scenario generation, when students issue natural language instructions in the learning environment (such as "I need to review math" or "I want to take a break"), the system will generate corresponding learning scenarios according to the instructions. For example: in the "review math" scenario, the system will automatically turn on the desk lights, turn off entertainment equipment, and recommend relevant learning materials; in the "rest" scenario, the system will automatically play relaxing music and remind students to relax their eyes; dynamic learning optimization, through the dynamic feedback component to perceive the student's learning status and environmental changes in real time, and dynamically optimize the learning scene. For example: if it is detected that the student is not paying attention, the system will remind the student to take a break or adjust the learning content; if the learning environment is not well lit, the system will automatically turn on the lights. The embodiment of the present invention realizes the personalized generation and dynamic optimization of the whole house smart scene through natural language instructions, multimodal data collection, deep semantic understanding and dynamic feedback mechanism. Whether in home, office, hotel, elderly care or education scenarios, the system can significantly improve the user experience, meet diverse needs, and have high flexibility and adaptability.

[0025] Embodiment 1: like Figure 1 As shown, an embodiment of the present invention provides a whole-house intelligent scene generation system based on natural language, comprising: The scene perception component is responsible for capturing the user's behavior habits (such as work and rest time, preferred temperature, light brightness, etc.) and environmental conditions (such as temperature, humidity, light intensity, etc.) in real time through multimodal data such as voice sensors, image sensors, and environmental sensors; and associating the user's instructions collected by the voice sensor with the behavior habits and environmental conditions; The semantic understanding component is responsible for using a large language model to perform deep semantic analysis on user voice commands. Based on the results of the deep semantic analysis, it obtains the associated behavioral habits and environmental status, and generates a scenario configuration plan based on the real-time status of all the equipment in the house. The dynamic feedback component is responsible for dynamically optimizing the scene generation strategy by continuously collecting user feedback and environmental data, and continuously adjusting the scene configuration based on user satisfaction, device operating status, and external environment changes.

[0026] In the above embodiments, the scene perception component builds a comprehensive user behavior and environment portrait through voice, image, environmental sensors, etc.; continuously captures user behavior patterns and preferences, and establishes a dynamic user model; establishes intelligent associations between voice commands and user habits and environmental status; provides a personalized service foundation: provides accurate data support for scene configuration; makes smart home more natural and more intimate. The semantic understanding component accurately understands the deep meaning of user commands; combines commands with user habits, environmental status, and device status; outputs the most optimized scene configuration suggestions; realizes natural interaction: allows users to control the home in the most natural way; enables the system to actively provide the best solution; and enables the home system to truly have the ability to "think". The dynamic feedback component continuously optimizes the scene strategy through user feedback, dynamically adjusts according to environmental changes and user satisfaction, and realizes continuous improvement of scene configuration; realizes system evolution, allowing the smart home system to continue to grow; provides personalized services, allowing the system to understand users more and more; improves system reliability, and ensures that the scene configuration is always optimal.

[0027] In summary, this embodiment realizes a complete closed loop from perception, understanding to optimization; it not only allows smart homes to understand users better, but also allows them to continuously learn and evolve, providing users with more and more considerate services. It represents the future development direction of smart homes and will greatly improve people's quality of life and happiness.

[0028] Embodiment 2: like Figure 2 As shown, based on Example 1, the scene perception component provided by the embodiment of the present invention includes: The speech decomposition module is responsible for extracting features from speech commands and decomposing them into feature vectors in multiple dimensions, such as time, space, and semantics. At the same time, the environmental state data (such as temperature, humidity, and light intensity) collected by the environmental sensor is mapped into an environmental state matrix in real time. The feature vector of the speech command is dynamically matched with the environmental state matrix through adaptive alignment to establish a preliminary association relationship. The graph construction module is responsible for constructing a behavior feature graph based on user behavior habit data (such as work and rest time, preferred temperature, light brightness, etc.), which contains multi-dimensional information such as time series, spatial distribution and intensity change; through behavior-command association; deeply couples the feature vector of the voice command with the behavior feature graph; The weight calculation module is responsible for real-time analysis of the correlation between the current environment status and historical behavior patterns. It assigns different weights to each feature dimension of the voice command through dynamic weight allocation. The weights are dynamically adjusted according to changes in the environment status and the evolution of behavioral habits. The vector generation module is responsible for multimodal data fusion of the feature vector of voice commands, the environmental state matrix and the behavioral feature map to generate an enhanced semantic vector that includes environmental adaptability and behavioral habits.

[0029] In the above embodiment, the voice decomposition module extracts features from the voice command and decomposes it into feature vectors of multiple dimensions such as time, space and semantics; it can accurately capture the core information of the voice command to ensure that the system can accurately understand the user's intention; at the same time, the environmental state data (such as temperature, humidity, light intensity, etc.) collected by the environmental sensor is mapped into an environmental state matrix in real time, providing basic data for subsequent matching and association. Significance: It can initially associate the voice command with the environmental state, thereby providing a basis for subsequent intelligent decision-making; through this association, it can better understand the user's needs in the current environment, thereby providing more personalized services. The graph construction module contains multi-dimensional information such as time series, spatial distribution and intensity changes; through the behavior-command association, the feature vector of the voice command is deeply coupled with the behavior feature map, further enhancing the system's understanding of user behavior. Significance: It can predict the user's needs based on the user's historical behavior habits and make corresponding adjustments in advance; the prediction ability not only improves the system's response speed, but also enhances the personalization of the user experience. The weight calculation module dynamically adjusts the weight according to the changes in the environmental state and the evolution of behavioral habits to ensure that the system can flexibly respond to various complex scenarios. Significance: The system can dynamically adjust the decision weight according to the actual situation, so as to make more reasonable judgments; the dynamic adjustment capability not only improves the intelligence level of the system, but also enhances its adaptability and robustness. The vector generation module enhances the semantic vector to more comprehensively reflect the real needs of users and provide a more accurate basis for the final decision of the system. Significance: Through multimodal data fusion, a more comprehensive and accurate semantic vector is generated, thereby improving the decision-making accuracy and intelligence level of the system; the enhanced semantic vector can not only better understand the needs of users, but also provide more personalized services according to environmental changes and the evolution of behavioral habits.

[0030] In summary, the modules of this embodiment together constitute an intelligent scene perception component, which can more accurately understand user needs and provide personalized services through multi-dimensional analysis of voice commands, environmental status and user behavior. The intelligent interactive system not only improves the user experience, but also enhances the adaptability and robustness of the system, providing strong technical support for future scenarios such as smart homes.

[0031] Embodiment 3: like Figure 3 As shown, based on Example 2, the speech decomposition module provided by the embodiment of the present invention includes: The multi-dimensional analysis submodule is responsible for decomposing the voice command into feature vectors of multiple dimensions such as time, space and semantics after feature extraction, which respectively represent the information of the voice command in different dimensions; environmental status data (such as temperature, humidity, light intensity, etc.) is collected through environmental sensors and mapped into an environmental status matrix, which contains the real-time status information of the current environment; The temporal dimension captures the temporal characteristics of voice commands, such as the duration and timing of the commands; The spatial dimension reflects the spatial characteristics of voice commands, such as the direction and location of the command source; The semantic dimension analyzes the semantic content of voice commands, such as the specific meaning of the command, keywords, etc. The hierarchical mapping submodule is responsible for dividing the environmental data into the basic layer, dynamic layer and associated layer through hierarchical mapping; The base layer contains the basic physical parameters of the environment (such as temperature, humidity, light intensity, etc.); The dynamic layer reflects the real-time changes in environmental conditions (such as temperature fluctuations, changing trends in light intensity, etc.); The correlation layer captures the correlation between environmental states (such as the correlation between temperature and humidity, the correspondence between light intensity and time, etc.); The dynamic matching submodule is responsible for preliminarily matching the temporal, spatial and semantic feature vectors of the voice command with the basic layer of the environment state matrix to find possible association points; through the analysis of the association layer, the semantic dimension of the voice command is deeply matched with the association information of the environment state matrix.

[0032] Among them, the time dimension feature extraction formula of the multi-dimensional analysis submodule is: ; In the formula, represents the time dimension feature vector, which represents the time characteristics of the voice command; A time domain waveform function representing a speech signal; Represents the frequency component function of the speech signal; and Indicates the start and end time points of the voice command; Indicates the time domain rate of change of the speech signal; Represents the logarithmic weighting of the frequency component, which is used to enhance the weight of high-frequency information; Spatial dimension feature extraction formula: ; In the formula, Represents the spatial dimension feature vector, indicating the source direction and location of the voice command; and Indicates the azimuth and elevation angles of the target direction; and Indicates i The azimuth and elevation angles of the microphones; Indicates i The weight coefficients of the microphones; Represents the standard deviation of the Gaussian distribution, which is used to smooth the spatial distribution; Indicates the number of microphones; Semantic dimension feature extraction formula: ; In the formula, Represents the semantic dimension feature vector, which represents the semantic content of the voice command; Indicates key words or phrases in voice commands; Represents word frequency-inverse document frequency, which is used to measure the frequency of keywords in documents The importance of Represents word embedding vector, which represents the semantic information of keywords; and Represents the weight coefficient, which is used to balance the contribution of different semantic features; Indicates the number of related documents; The basic layer mapping formula of the hierarchical mapping submodule is: ; In the formula, Represents the base layer matrix, which contains the basic physical parameters of the environment; represents the temperature vector, ; represents the humidity vector, ; represents the light intensity vector, ; n Indicates the number of sensors; Dynamic layer mapping formula: ; In the formula, Represents a dynamic layer matrix, reflecting real-time changes in environmental status; Represents the base layer matrix at time Status; represents the time rate of change of the base layer matrix; Indicates the current time point; Represents the time attenuation coefficient, which is used to smooth dynamic changes; Association layer mapping formula: ; In the formula, represents the correlation layer matrix, capturing the correlation between environmental states; and Represents the first and parameters; Indicates and The correlation coefficient between the parameters; Indicates the number of base layer parameters; The preliminary matching formula of the dynamic matching submodule is: ; In the formula, represents the preliminary matching score, which indicates the correlation between the voice command feature vector and the environment state matrix; Temporal, spatial or semantic feature vectors representing speech instructions; Base layer parameters representing the environment state matrix; Represents the weight coefficient, which is used to adjust the matching contribution of different features; represents the number of eigenvectors; Depth matching formula: ; In the formula, represents the deep matching score, which indicates the correlation between the semantic dimension of the voice command and the environment state matrix; A semantic feature vector representing the voice command; Represents the associated layer parameters of the environment state matrix; Represents the weight coefficient, which is used to adjust the matching contribution of different semantic features; and Represents the number of semantic feature vectors and associated layer parameters; the above design fully considers the complexity of the speech decomposition module and the actual application needs, and realizes efficient collaborative analysis of voice commands and environmental status through multi-dimensional analysis, hierarchical mapping and dynamic matching.

[0033] In the above embodiment, the multi-dimensional parsing submodule helps the system understand the timeliness of the instruction, such as time-related semantics such as "now" or "ten minutes later"; reflects the spatial characteristics of the voice instruction, such as the direction and location of the instruction source, and can identify the spatial background of the instruction, such as "the light on the left" or "the temperature of the room"; analyzes the semantic content of the voice instruction, such as the specific meaning and keywords of the instruction, and can accurately understand the user's intention. Significance: Through multi-dimensional parsing, the context information of the voice instruction can be more comprehensively understood, thereby improving the accuracy and intelligence level of voice recognition. For example, combining time, space and semantic information, it can respond to user needs more accurately, such as "turn up the light on the left now" or "lower the temperature of the room in ten minutes". Hierarchical mapping submodule, the basic physical parameters of the environment (such as temperature, humidity, light intensity, etc.) are the basis of environmental perception, providing raw data support for dynamic analysis and association analysis; through the dynamic layer, the system can capture the real-time changes in the environmental state, so as to make a more timely response; the association layer provides more complex context information to help it understand the interaction between environmental states. Significance: Hierarchical mapping enables the system to analyze environmental data step by step from basic to complex, so as to have a more comprehensive understanding of the environmental status. It not only improves the system's perception ability, but also provides richer data support for voice command matching. The dynamic matching submodule deeply matches the semantic dimension of the voice command with the correlation information of the environmental status matrix through the analysis of the association layer; for example, the humidity setting of the air conditioner may be adjusted to meet the needs of the user in combination with the correlation between temperature and humidity. Significance: Dynamic matching can make more intelligent responses based on the context of the voice command and the real-time changes of the environmental status; it not only improves the user experience, but also can better adapt to complex environmental changes.

[0034] In summary, the speech decomposition module of this embodiment achieves deep integration of speech commands and environmental status data through multi-dimensional analysis, hierarchical mapping and dynamic matching; improves the accuracy and intelligence level of speech recognition; enhances the system's ability to perceive environmental status; achieves dynamic matching between speech commands and environmental status, and can respond to user needs more intelligently. Its significance lies in: providing users with a more natural and intelligent voice interaction experience; being able to better adapt to complex environmental changes and improve the overall intelligence level; providing a technical foundation for future scenarios such as smart homes, and promoting the further development of human-computer interaction.

[0035] Embodiment 4: like Figure 4 As shown, based on Example 2, the graph construction module provided in this embodiment of the present invention includes: The feature spatiotemporal decomposition submodule is responsible for decomposing the user behavior habit data into multiple time series segments according to the time dimension to capture the time characteristics of the user behavior; for example, analyzing the user's daily routine and device usage frequency; decomposing the user behavior habit data into multiple spatial distribution areas according to the spatial dimension to capture the spatial characteristics of the user behavior; for example, analyzing the user's activity frequency in different rooms and the device usage location; decomposing the user behavior habit data into multiple intensity change curves according to the intensity dimension to capture the intensity characteristics of the user behavior; for example, analyzing the changes in the user's preference for temperature and light intensity; The feature hierarchical modeling submodule is responsible for building a basic feature model of user behavior based on time series, spatial distribution and intensity change data; for example, forming a baseline of user behavior in different time periods and different spatial areas; capturing real-time changes and trends in user behavior to form a dynamic feature model; for example, analyzing the changing patterns of user behavior with environmental factors such as seasons and weather; analyzing the correlation between user behaviors to form a correlation feature model, for example, analyzing the correlation between user schedules and device usage frequency, or the synergistic relationship between temperature preference and light intensity; The deep coupling submodule is responsible for mapping the feature vectors of voice commands (such as time, space, and semantic dimensions) with the behavioral feature map to find potential associations between commands and behaviors; for example, matching the "raise the temperature" command with the user's temperature preferences in different time periods; identifying user behavior patterns under the current environmental state through the hierarchical model of the behavioral feature map, and deeply coupling them with voice commands; for example, predicting the user's temperature preference at a specific time point based on the user's historical schedule.

[0036] Among them, the time dimension feature extraction formula of the feature spatiotemporal decomposition submodule is: ; In the formula, Represents the time dimension feature vector, which represents the time characteristics of user behavior; A time domain waveform function representing user behavior habit data; A frequency component function representing user behavior habit data; , Indicates the start and end time points of user behavior; Indicates the time domain change rate of user behavior habit data; Represents the logarithmic weighting of the frequency component, which is used to enhance the weight of high-frequency information; Indicates the current time point; The Gaussian smoothing coefficient representing the time dimension is used to capture the local characteristics of the time segment; Spatial dimension feature extraction formula: ; In the formula, Represents the spatial dimension feature vector, which represents the spatial distribution characteristics of user behavior; , Indicates the azimuth and elevation angles of the target position; , Indicates The central azimuth and elevation angles of a spatial region; Indicates The weight coefficient of each spatial region; The Gaussian smoothing coefficient representing the spatial dimension is used to capture the local characteristics of the spatial distribution; Indicates The intensity of activity in a spatial area; Indicates the number of spatial regions; Intensity dimension feature extraction formula: ; In the formula, Represents the intensity dimension feature vector, which represents the intensity characteristics of user behavior; Indicates the intensity change curve of user behavior habit data; A frequency component function representing intensity variations; Indicates the time domain change rate of the intensity change curve; Logarithmic weighting of frequency components representing intensity variations; The Gaussian smoothing coefficient representing the intensity dimension is used to capture the local characteristics of intensity variations; The basic feature model formula of the feature hierarchical modeling submodule is: ; In the formula, Represents the basic feature model matrix, which contains the basic features of user behavior; Represents the time dimension feature vector; Represents the spatial dimension feature vector; represents the intensity dimension feature vector; Dynamic feature model formula: ; In the formula, Represents the dynamic feature model matrix, reflecting the real-time changes in user behavior; Represents the state of the basic feature model matrix at time t; Represents the time rate of change of the basic feature model matrix; Indicates the current time point; Represents the time attenuation coefficient, which is used to smooth dynamic changes; Indicates The influence function of environmental factors (such as season and weather); Indicates The weight coefficient of each environmental factor; Indicates the number of environmental factors; Association feature model formula: ; In the formula, Represents the correlation feature model matrix, capturing the correlation between user behaviors; , Represents the first and parameters; Indicates and The correlation coefficient between the parameters; The Gaussian smoothing coefficient representing the correlation dimension is used to capture the local characteristics of the correlation; represents the number of basic feature model parameters; The instruction-behavior mapping formula of the deep coupling sub-module is: ; In the formula, represents the command-behavior mapping matrix, which represents the potential association between the speech command feature vector and the behavior feature map; A feature vector representing the speech command (temporal, spatial, or semantic dimension); Base layer parameters representing the behavioral feature map; Represents the weight coefficient, which is used to adjust the mapping contribution of different features; The Gaussian smoothing coefficient representing the mapping dimension is used to capture the local characteristics of the mapping relationship; Hierarchical coupling formula: ; In the formula, represents a hierarchical coupling matrix, which indicates the deep coupling between the speech command feature vector and the dynamic feature model and the associated feature model; , A feature vector representing the speech command (temporal, spatial, or semantic dimension); represents the first parameters; Represents the first q parameters; , Represents the weight coefficient, which is used to adjust the coupling contribution of different features; The Gaussian smoothing coefficient representing the coupling dimension is used to capture the local characteristics of the coupling relationship; , , , Represents the number of feature vectors, dynamic feature model parameters, and associated feature model parameters. The above formula fully reflects the functions of "feature spatiotemporal decomposition", "feature hierarchical modeling", and "deep coupling", especially the deep coupling part, which realizes the efficient collaborative analysis of voice commands and user behavior feature maps through command-behavior mapping and hierarchical coupling.

[0037] In the above embodiments, the feature spatiotemporal decomposition submodule decomposes the user behavior data into multiple time series segments according to the time dimension, which can capture the time characteristics of the user behavior; for example, analyze the user's daily routine, device usage frequency, etc.; the spatial dimension decomposition decomposes the user behavior data into multiple spatial distribution areas according to the spatial dimension, which can capture the spatial characteristics of the user behavior; for example, analyze the user's activity frequency in different rooms, the location of device use, etc.; the intensity dimension decomposition decomposes the user behavior data into multiple intensity change curves according to the intensity dimension, which can capture the intensity characteristics of the user behavior; for example, analyze the changes in the user's preference for temperature and light intensity. Significance: Through multi-dimensional decomposition, it is possible to more accurately understand the user's behavioral habits and preferences, provide data support for personalized services, help better adapt to the user's living environment, and provide services that are more in line with user needs. The feature hierarchical modeling submodule builds a basic feature model of user behavior based on time series, spatial distribution and intensity change data; for example, it forms a baseline of user behavior in different time periods and different spatial areas; it captures real-time changes and trends in user behavior to form a dynamic feature model; for example, it analyzes how user behavior changes with environmental factors such as seasons and weather; the associated feature model analyzes the correlation between user behaviors to form an associated feature model; for example, it analyzes the correlation between user schedules and device usage frequency, or the synergistic relationship between temperature preference and light intensity. Significance: Through the dynamic feature model, it is possible to predict future user behavior trends, adjust system settings in advance, and provide more intelligent services; through the associated feature model, it is possible to discover potential connections between user behaviors, optimize system response strategies, and improve user experience. The deep coupling submodule maps the feature vectors of voice commands (such as time, space, and semantic dimensions) with the behavioral feature map to find the potential correlation between commands and behaviors; for example, matching the "turn up the temperature" command with the user's temperature preferences in different time periods; through the hierarchical model of the behavioral feature map, identifying the user's behavior patterns in the current environment state, and deeply coupling with the voice command; for example, combining the user's historical schedule to predict the user's temperature preference at a specific time point. Significance: Through deep coupling, it can more intelligently understand and respond to the user's voice commands and provide more personalized services; through accurate behavioral pattern recognition and command matching, it can significantly improve the user's experience and reduce the user's operational burden.

[0038] In summary, the graph construction module of this embodiment achieves accurate analysis and intelligent response to user behavior through multi-dimensional decomposition, hierarchical modeling and deep coupling. It not only improves the intelligence level of the system, but also provides users with a more personalized and considerate service experience. Through these technical means, we can better understand and adapt to user needs and realize true smart home and personalized services.

[0039] Embodiment 5: like Figure 5 As shown, based on Example 2, the weight calculation module provided in this embodiment of the present invention includes: The data mapping submodule is responsible for collecting the current environmental status data (such as temperature, humidity, light intensity, etc.) through environmental sensors and mapping it into an environmental status matrix. At the same time, the voice decomposition module extracts features from the user's voice commands and generates feature vectors containing multiple dimensions such as time, space, and semantics. The feature recognition submodule is responsible for analyzing the correlation between the current environment state and the user behavior pattern by combining the historical behavior feature map; by comparing the current environment state with the historical behavior pattern, it identifies which feature dimensions are more important in the current scenario; The weight assignment submodule is responsible for dynamically assigning weights to each feature dimension of the voice command.

[0040] In the above embodiment, the data mapping submodule converts the environmental state and voice instructions into computable structured data, providing a basis for analysis and matching. Significance achieved: The data mapping submodule realizes the digital expression of environmental state and voice instructions, enabling the system to process multimodal data (environmental data and voice data) in a unified framework; mapping not only provides data support for subsequent feature recognition and weight allocation, but also ensures the system's real-time perception of environmental changes, laying the foundation for dynamic response. The feature recognition submodule combines the historical behavior feature map to analyze the correlation between the current environmental state and the user's behavior pattern; by comparing the current environmental state with the historical behavior pattern, it identifies which feature dimensions are more important in the current scene; for example, in a high temperature environment, temperature-related feature dimensions may be identified as more important; and at night, light-related feature dimensions may be given priority. Significance achieved: The feature recognition submodule achieves a deep understanding of user intentions by mining the correlation between environmental state and user behavior patterns; not only based on the current scene data, but also combined with the user's long-term behavior habits, so that the system's response is more in line with the user's actual needs. The intelligence level of the system has been improved, enabling it to extract key features from massive data and provide a basis for weight allocation. The weight allocation submodule dynamically allocates weights to each feature dimension of the voice command based on the analysis results of the feature recognition submodule; for example, if the current ambient temperature is high and the user's historical behavior shows that they are more inclined to adjust the air-conditioning temperature in a high temperature environment, then the temperature-related feature dimensions will be given a higher weight; the weight allocation is dynamic and will be adjusted continuously with changes in environmental conditions and the evolution of user behavior habits. Significance achieved: The weight allocation submodule achieves accurate interpretation and response to voice commands by dynamically adjusting weights; the dynamic allocation mechanism not only takes into account the current environmental conditions, but also combines the user's behavior habits, making the system's response more personalized and intelligent. The adaptability and user experience of the system have been improved, enabling it to provide the best interaction solution based on different scenarios and user needs.

[0041] In summary, the weight calculation module of this embodiment realizes in-depth analysis and dynamic matching of environmental status, user voice commands and historical behavior patterns through the coordinated work of three submodules: data mapping, feature recognition and weight allocation. The multi-level analysis and matching mechanism not only improves the system's environmental adaptability and user behavior perception capabilities, but also significantly enhances the naturalness and intelligence of voice interaction. It provides users with a more accurate, personalized and intelligent voice interaction experience, and promotes the further development of human-computer interaction technology.

[0042] Embodiment 6: like Figure 6 As shown, based on Example 1, the semantic understanding component provided by the embodiment of the present invention includes: The model generation module is responsible for using semantic analysis mechanism to convert the user's natural language instructions into executable scene configuration solutions; through deep semantic association, combined with the real-time status of all the equipment in the house, a dynamic scene model is generated; for example, when the user says "I want to watch a movie", it will not only turn on the TV, but also automatically adjust the curtains and lights according to the current light conditions, and recommend suitable movies according to the user's viewing habits; The model triggering module is responsible for receiving voice commands and triggering the large language model; The solution execution module is responsible for generating and executing highly personalized scenario solutions through perception and understanding mechanisms; for example, when the user switches from the "preparing dinner" to the "dining" scene, the system will automatically adjust the operating status of the kitchen equipment (such as turning off the oven and turning on the dining table lighting) and play background music according to the user's dining habits.

[0043] In the above embodiment, the semantic analysis mechanism of the model generation module can understand the user's natural language instructions and convert them into executable scene configuration schemes; through deep semantic association, the module can combine the real-time status of the whole house equipment to generate a dynamic scene model; for example, when the user says "I want to watch a movie", the module will not only turn on the TV, but also automatically adjust the curtains and lights according to the current light conditions, and recommend suitable movies according to the user's viewing habits. Significance: The "intelligence" and "personalization" of the smart home system are realized; it is not just a simple response to user instructions, but through a deep understanding of the user's intentions and environmental conditions, it provides a scene configuration that is more in line with user needs; it not only improves the user experience, but also reduces the user's operating burden, making the smart home system more "understanding you". The model trigger module receives voice instructions and triggers the large language model for processing; it can quickly recognize the user's voice input and convert it into instructions that the system can understand, and then start the subsequent scene configuration and execution process. Significance: The "instant response" of the smart home system is realized; through efficient voice recognition and instruction triggering mechanism, users can quickly start complex scene configurations through simple voice instructions, greatly improving the system's response speed and user experience; at the same time, this also lays the foundation for personalized scene execution. The solution execution module generates and executes highly personalized scenario solutions through the perception and understanding mechanism; for example, when the user switches from the "preparing dinner" to the "dining" scene, the system automatically adjusts the operating status of the kitchen equipment (such as turning off the oven and turning on the table lighting) and plays background music according to the user's dining habits. Significance: It realizes the "seamless switching" and "personalized execution" of the smart home system; it can automatically adjust the device status according to the user's behavioral habits and environmental changes to provide the most appropriate scenario experience. It not only improves the user's quality of life, but also makes the smart home system closer to the user's daily living habits, truly realizing the perfect integration of "intelligence" and "home".

[0044] In summary, the various modules of the semantic understanding component of this embodiment work together to achieve a high degree of intelligence and personalization of the smart home system. The model generation module generates dynamic scene configurations through deep semantic analysis, the model trigger module quickly starts the system through voice commands, and the solution execution module provides a seamless switching scene experience through personalized execution. This enables the smart home system to not only respond to user commands, but also understand user needs and provide more intimate and intelligent services.

[0045] Embodiment 7: like Figure 7 As shown, based on Example 6, the model generation module provided in this embodiment of the present invention includes: The multimodal data processing submodule is responsible for converting the user's natural language instructions into structured semantic information, extracting key actions, objects and context information in the instructions; capturing the user's behavior patterns (such as entering the kitchen, opening the refrigerator, etc.) in real time and converting them into time series features; capturing the current environmental state in real time and quantifying it into a computable feature vector; combining the user's historical behavior data (such as common recipes, cooking habits) and environmental change trends to form a multi-dimensional data foundation; The semantic association and demand inference submodule is responsible for deeply associating voice commands with user behavior patterns to determine whether the behavior conforms to the typical pattern of commands (such as the association between the "prepare dinner" command and behaviors such as taking ingredients and turning on the oven); inferring changes in user needs by analyzing the association between user behavior patterns and environmental changes (such as frequent walking and rising temperatures may indicate that the user is cooking); and optimizing the execution effect of commands by associating voice commands with environmental changes (such as the association between the "prepare dinner" command and light and temperature to automatically adjust lighting or air conditioning); The dynamic scene model generation submodule is responsible for generating a state model of the current scene based on multimodal data (voice commands, behavior patterns, environmental status), including the user's behavior status, device operation status, and environmental parameters; integrating historical behavior data and environmental change trends into the scene model to form a deep understanding of user habits and preferences; and predicting possible changes in user needs (such as the transition from "preparing dinner" to "dining" scenes) by analyzing the current scene status and historical data, and generating corresponding prediction information.

[0046] In the above embodiment, the multimodal data processing submodule realizes accurate understanding of user instructions by converting the user's natural language instructions into structured semantic information, extracting key actions, objects and context information; at the same time, it can capture the user's behavior pattern and environmental status in real time, and quantify them into computable feature vectors, combined with historical behavior data, to form a multi-dimensional data foundation. Significance: It provides high-quality data support for semantic association, demand inference and scenario generation; enables the system to understand the user's intention more accurately, and lays a data foundation for personalized services; through the fusion of multimodal data, it can more comprehensively perceive the user's needs and environmental changes, thereby providing more intelligent services. The semantic association and demand inference submodule deeply associates voice instructions with the user's behavior pattern, and can determine whether the behavior conforms to the typical pattern of the instruction. At the same time, by analyzing the association between user behavior and environmental changes, the user's demand changes are inferred, and the execution effect of the instruction is optimized. Significance: It enables a more intelligent understanding of the user's real needs, rather than just mechanically executing instructions; for example, when the user says "prepare dinner", it can not only perform related operations, but also automatically adjust lighting, temperature and other parameters according to the user's behavior pattern and environmental changes to provide more considerate services; this intelligent demand inference capability greatly improves the user experience. The dynamic scene model generation submodule generates a state model of the current scene based on multimodal data, including the user's behavior state, device operation state and environmental parameters; at the same time, it integrates historical behavior data and environmental change trends into the scene model to form a deep understanding of user habits and preferences, and predict possible changes in user needs. Significance: It can dynamically generate scene models to reflect user needs and environmental changes in real time. By predicting possible changes in user needs, the system can prepare in advance, such as automatically adjusting lighting and music in the transition from "prepare dinner" to "dining" scenes to provide users with a seamless experience; this dynamic scene generation capability makes the system more intelligent and humane.

[0047] In summary, the submodules of this embodiment together constitute an intelligent scene generation and configuration system, which can provide users with highly intelligent and personalized services through multimodal data processing, semantic association, demand inference and dynamic scene generation. The significance lies in: providing seamless intelligent services through accurate understanding of user needs; enabling the system to perceive and respond to user needs in real time through multimodal data fusion and dynamic scene generation; generating highly personalized scene configuration solutions through historical behavior data and preference analysis to meet the personalized needs of users.

[0048] Embodiment 8: like Figure 8 As shown, based on Example 7, the dynamic scene model generation submodule provided in this embodiment of the present invention includes: The association analysis unit is responsible for analyzing the association between the user's behavior pattern and environmental changes. For example, if the user moves frequently and the ambient temperature rises, it is inferred that the user may be cooking. The association analysis is based on the current behavior and environmental data, combined with historical data, to form a deep understanding of the user's habits. The data integration unit is responsible for generating a state model of the current scene based on the collected multimodal data, including the user's behavior status (such as "getting ingredients"), equipment operation status (such as "preheating the oven"), and environmental parameters (such as "the kitchen temperature rises"). By integrating historical behavior data and environmental change trends into the model, the system can more accurately capture changes in user needs; The result prediction unit is responsible for predicting possible changes in user needs by analyzing the current scene status and historical data. For example, when the user completes the "preparing dinner" behavior sequence, the system predicts that the user is about to enter the "dining" scene. The prediction is based on the current behavior and environmental data, and the user's historical habits (such as meal times, commonly used tableware, etc.). It generates corresponding prediction information and adjusts environmental parameters in advance (such as adjusting lighting and lowering air-conditioning temperature) to adapt to changes in user needs.

[0049] In the above embodiment, the association analysis unit* can infer the user's current activity status (such as cooking, dining, etc.) by analyzing the association between the user's behavior pattern (such as frequent walking, turning on the device) and environmental changes (such as temperature increase, light changes); based on the current behavior and environmental data, combined with historical data, the system can capture the user's behavior trends in real time and dynamically adjust the results of the association analysis; through long-term accumulation of historical data, it can form a deep understanding of user habits, such as the user's typical behavior pattern in a specific time period. Significance achieved: Through the association analysis of behavior and environment, the user's current activity status can be judged more accurately to avoid misjudgment or delayed response; the user's needs can be actively inferred based on the user's behavior pattern and environmental changes, rather than relying solely on the user's explicit instructions; the results of the association analysis provide high-quality data input for the data integration unit and the result prediction unit to ensure the accuracy and reliability of the prediction. The data integration unit* integrates multi-dimensional data such as user behavior status (such as "taking ingredients"), equipment operation status (such as "preheating the oven"), and environmental parameters (such as "the temperature in the kitchen rises") to generate a state model of the current scene; by integrating historical behavior data and environmental change trends into the model, it can more accurately capture changes in user needs and avoid dependence on a single data source; it can dynamically update the scene model based on real-time collected data to ensure that the model always reflects the current real state. Significance achieved: Through the fusion of multimodal data, a comprehensive model including user behavior, equipment status, and environmental parameters can be generated, providing a basis for subsequent prediction and optimization; the integration of historical data enables the system to better understand user habits and preferences, thereby more accurately capturing changes in user needs; the dynamically updated scene model enables the system to quickly adapt to environmental changes and changes in user behavior, ensuring the timeliness and accuracy of the response. The result prediction unit can predict possible changes in user needs (such as the transition from "preparing dinner" to "dining" scenes) by analyzing the current scene status and historical data; generate corresponding prediction information based on the prediction results (such as adjusting lighting, lowering air conditioning temperature), and continuously optimize the prediction logic to improve accuracy; and can adjust environmental parameters (such as light, temperature, and device status) in advance according to the prediction results to adapt to changes in user needs. Significance achieved: By predicting changes in user needs, it can achieve intelligent transition of scenes (such as from "preparing dinner" to "dining") and improve user experience; it can actively adjust environmental parameters based on prediction results without manual intervention by users, reflecting the core value of intelligent services; by combining users' historical habits (such as meal times, commonly used tableware, etc.), it can provide more personalized services to meet users' unique needs.

[0050] In summary, this embodiment improves the accuracy of scene understanding through the correlation analysis of behavior and environment, providing a data basis for prediction; through the deep integration of multimodal data, a comprehensive scene state model is constructed to improve the accuracy of demand capture; through accurate prediction of demand changes, intelligent scene transition and active service are realized to improve user experience. The three units work together to achieve a deep understanding and intelligent response to user behavior and demand changes, and ultimately achieve the goal of improving user experience and enhancing the intelligence level of the system.

[0051] Fig. 9 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present invention is shown.

[0052] The electronic device may include a central processing unit / microprocessor / main control chip, etc. 1; a storage medium 2, coupled to the central processing unit / microprocessor / main control chip, etc. 1, and storing computer executable instructions therein, for performing the steps of each method of an embodiment of the present invention when executed by the processor.

[0053] The central processing unit / microprocessor / main control chip etc. 1 may include but is not limited to, for example, one or more processors or microprocessors etc.

[0054] The storage medium 2 may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disk, floppy disk, solid-state drive, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, etc.).

[0055] In addition, the electronic device may also include (but not limited to) a data bus 3, an input / output bus / external bus / device bus 4, a display 5, and input / output devices 6 (eg, keyboard, mouse, speaker, etc.).

[0056] The central processing unit / microprocessor / main control chip etc. 1 can communicate with external devices ( 5 , 6 etc.) through the I / O bus 4 via a wired or wireless network (not shown).

[0057] The storage medium 2 may also store at least one computer executable instruction for executing the various functions and / or method steps in the embodiments described in the present technology when the instruction is executed by the central processing unit / microprocessor / main control chip 1.

[0058] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.

[0059] Fig.10 A schematic diagram of a computer-readable storage medium according to an embodiment of the present invention is shown.

[0060] like Fig.10 As shown, instructions are stored on the non-transitory computer-readable storage medium 8, and the instructions are, for example, computer-readable instructions 7. When the computer-readable instructions 7 are executed by the processor, the various methods described above can be executed. The non-transitory computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-transitory non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc. For example, the non-transitory computer-readable storage medium 8 can be connected to a computing device such as a computer, and then, when the computing device runs the computer-readable instructions 7 stored on the computer-readable storage medium 8, the various methods described above can be performed.

[0061] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0062] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0063] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0064] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions for executing all or part of the steps of the various embodiments of the method of the present invention through a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name in English: Read-Only Memory, English abbreviation: ROM), random access memory (full name in English: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program code.

[0065] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A whole-house intelligent scene generation system based on natural language, characterized in that: Include: The scene perception component is responsible for capturing the user's behavior habits and environmental status in real time through multimodal data from voice sensors, image sensors, and environmental sensors; and associating the user's instructions collected by the voice sensor with the behavior habits and environmental status; The semantic understanding component is responsible for using a large language model to perform deep semantic analysis on user voice commands. Based on the results of the deep semantic analysis, it obtains the associated behavioral habits and environmental status, and generates a scenario configuration plan based on the real-time status of all the equipment in the house. The dynamic feedback component is responsible for dynamically optimizing the scene generation strategy by continuously collecting user feedback and environmental data, and continuously adjusting the scene configuration based on user satisfaction, device operating status, and external environment changes.

2. The natural language-based whole-house intelligent scene generation system according to claim 1, characterized in that: Scene perception components, including: The speech decomposition module is responsible for extracting features from speech commands and decomposing them into feature vectors in multiple dimensions, including time, space, and semantics. At the same time, the environmental state data collected by the environmental sensor is mapped into an environmental state matrix in real time. The feature vector of the speech command is dynamically matched with the environmental state matrix through adaptive alignment to establish a preliminary association relationship. The graph construction module is responsible for constructing a behavior feature graph based on user behavior habit data, including multi-dimensional information such as time series, spatial distribution, and intensity change; through behavior-command association, the feature vector of the voice command is deeply coupled with the behavior feature graph; The weight calculation module is responsible for real-time analysis of the correlation between the current environment status and historical behavior patterns. It assigns different weights to each feature dimension of the voice command through dynamic weight allocation. The weights are dynamically adjusted according to changes in the environment status and the evolution of behavioral habits. The vector generation module is responsible for multimodal data fusion of the feature vector of voice commands, the environmental state matrix and the behavioral feature map to generate an enhanced semantic vector that includes environmental adaptability and behavioral habits.

3. The natural language-based whole-house intelligent scene generation system according to claim 2, characterized in that: Speech decomposition module, including: The multi-dimensional analysis submodule is responsible for decomposing the voice command into feature vectors of time, space and semantic dimensions after feature extraction, which respectively represent the information of the voice command in different dimensions; the environmental status data is collected through environmental sensors and mapped into an environmental status matrix, which contains the real-time status information of the current environment; The hierarchical mapping submodule is responsible for dividing the environmental data into the basic layer, dynamic layer and associated layer through hierarchical mapping; The dynamic matching submodule is responsible for preliminarily matching the temporal, spatial and semantic feature vectors of the voice command with the basic layer of the environment state matrix to find possible association points; through the analysis of the association layer, the semantic dimension of the voice command is deeply matched with the association information of the environment state matrix.

4. The natural language-based whole-house intelligent scene generation system according to claim 3, characterized in that: The time dimension of the multi-dimensional analysis submodule captures the temporal characteristics of voice commands; the spatial dimension reflects the spatial characteristics of voice commands; and the semantic dimension analyzes the semantic content of voice commands.

5. The natural language-based whole-house intelligent scene generation system according to claim 3, characterized in that: The base layer of the hierarchical mapping submodule contains the basic physical parameters of the environment; the dynamic layer reflects the real-time changes of the environment state; and the correlation layer captures the correlation between the environment states.

6. The natural language-based whole-house intelligent scene generation system according to claim 2, characterized in that: Graph building modules, including: The feature spatiotemporal decomposition submodule is responsible for decomposing the user behavior habit data into multiple time series segments according to the time dimension to capture the temporal characteristics of the user behavior; decomposing the user behavior habit data into multiple spatial distribution areas according to the spatial dimension to capture the spatial characteristics of the user behavior; decomposing the user behavior habit data into multiple intensity change curves according to the intensity dimension to capture the intensity characteristics of the user behavior; The feature hierarchical modeling submodule is responsible for building a basic feature model of user behavior based on time series, spatial distribution and intensity change data; capturing real-time changes and trends in user behavior to form a dynamic feature model; analyzing the correlation between user behaviors to form a correlation feature model. The deep coupling submodule is responsible for mapping the feature vectors of voice commands with the behavioral feature map to find the potential correlation between commands and behaviors. Through the hierarchical model of the behavioral feature map, it identifies the user behavior pattern under the current environmental state and deeply couples it with the voice commands.

7. The natural language-based whole-house intelligent scene generation system according to claim 2, characterized in that: Weight calculation module, including: The data mapping submodule is responsible for collecting the current environmental status data through environmental sensors and mapping it into an environmental status matrix. At the same time, the voice decomposition module extracts features from the user's voice commands and generates feature vectors containing multiple dimensions of time, space and semantics. The feature recognition submodule is responsible for analyzing the correlation between the current environment state and the user behavior pattern by combining the historical behavior feature map; by comparing the current environment state with the historical behavior pattern, it identifies which feature dimensions are more important in the current scenario; The weight assignment submodule is responsible for dynamically assigning weights to each feature dimension of the voice command.

8. The natural language-based whole-house intelligent scene generation system according to claim 1, characterized in that: Semantic understanding components, including: The model generation module is responsible for using a semantic analysis mechanism to convert the user's natural language instructions into executable scene configuration solutions; through deep semantic association, combined with the real-time status of all the equipment in the house, a dynamic scene model is generated; The model triggering module is responsible for receiving voice commands and triggering the large language model; The solution execution module is responsible for generating and executing highly personalized scenario solutions through perception and understanding mechanisms.

9. The natural language-based whole-house intelligent scene generation system according to claim 8, characterized in that: Model generation module, including: The multimodal data processing submodule is responsible for converting the user's natural language instructions into structured semantic information, extracting key actions, objects, and context information in the instructions; capturing the user's behavior patterns in real time and converting them into time series features; Capture the current environment status in real time and quantify it into a computable feature vector; combine the user's historical behavior data and environmental change trends to form a multi-dimensional data foundation; The semantic association and demand inference submodule is responsible for deeply associating voice commands with user behavior patterns to determine whether the behavior conforms to the typical pattern of commands; inferring changes in user needs by analyzing the association between user behavior patterns and environmental changes; and analyzing the association between voice commands and environmental changes to optimize the execution effect of commands; The dynamic scene model generation submodule is responsible for generating a state model of the current scene based on multimodal data, including the user's behavior status, device operation status, and environmental parameters; integrating historical behavior data and environmental change trends into the scene model to form a deep understanding of user habits and preferences; By analyzing the current scene status and historical data, possible changes in user demand are predicted and corresponding forecast information is generated.

10. The natural language-based whole-house intelligent scene generation system according to claim 9, characterized in that: Dynamic scene model generation submodule, including: The association analysis unit is responsible for analyzing the association between user behavior patterns and environmental changes. The association analysis is based on current behavior and environmental data, combined with historical data to form a deep understanding of user habits. The data integration unit is responsible for generating a state model of the current scene based on the collected multimodal data, including the user's behavior status, device operation status, and environmental parameters; by integrating historical behavior data and environmental change trends into the model; The result prediction unit is responsible for predicting possible changes in user needs by analyzing the current scene status and historical data; the prediction is based on current behavior and environmental data, and the user's historical habits; it generates corresponding prediction information and adjusts environmental parameters in advance to adapt to changes in user needs.

Citation Information

Patent Citations

  • Man-machine interaction system in smart home space and application method thereof

    CN115291718A

  • High-risk user complaint early warning method based on fusion algorithm

    CN115828167A

  • Voice remote control method and remote control system

    CN118942455A

  • Digital interaction enhancement system based on multi-mode voice

    CN119673154A

  • Video timestamp event identification and reasoning method based on multi-modal large model

    CN119723431A

Cited By

  • Multi-information fusion voice interaction method and system for range hood

    CN120544569A

  • Scene adaptive recommendation method in smart home based on multi-modal data

    CN121256148A

  • Scene-Adaptive Recommendation Methods for Smart Homes Based on Multimodal Data

    CN121256148B

  • State information query method and device of smart home equipment

    CN121301417A

  • Self-adaptive display adjusting system based on multi-mode perception

    CN121327752A