Intelligent auxiliary image-text generation multi-home device interaction method and storage medium
By combining cellular automata and LSTM models, more personalized and coherent interactive text for home appliances is generated, solving the problems of insufficient creativity and contextual understanding in traditional technologies and improving the user experience.
Patent Information
- Application Number
- CN202411632009.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Traditional intelligent assisted image and text generation technologies suffer from a lack of creativity and personalization, limited flexibility, and insufficient contextual understanding, resulting in a poor user experience.
We employ a cellular automata model for interaction prediction, combine it with LSTM to generate interactive text, and integrate the outputs of both models by fusing weights to generate more personalized, coherent, and context-aware interactive text.
It enhances the personalization and consistency of home appliance interaction, strengthens the emotional connection between users and devices, and improves user satisfaction and the smoothness and relevance of the interactive experience.
Smart Images

Figure CN119739045B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of smart home, and in particular to a multi-home device interaction method for smart auxiliary image-text generation and a storage medium. BACKGROUND
[0002] The interaction between home appliances is called home device interaction. In recent years, thanks to the rapid development of Internet of Things technology, home device interaction has entered an era of higher intelligence and greater convenience. Internet of Things technology enables various appliances in the home (such as smart lighting, smart speakers, smart door locks, etc.) to be connected to each other through the Internet, thereby enabling remote monitoring, control, and data transmission, and further enabling deeper intelligent interaction between these appliances.
[0003] Voice recognition technology enables people to communicate with home devices by issuing voice commands. This technology is known as voice assistants (such as Siri, Alexa, Google Assistant, etc.). These voice assistants can receive user instructions and then perform related actions. To achieve this goal, they need the help of natural language processing technology, which is a technology specifically designed to analyze and understand user input text to help appliances generate natural language responses or instructions, thereby achieving a more intelligent working state. In addition, if we install various sensors (such as temperature sensors, light sensors, etc.) on home devices, these sensors can detect changes in the environment and interact accordingly based on changes in the environment.
[0004] Traditional intelligent auxiliary image-text generation technology is based on predefined rules and conditions. The generator can generate corresponding image-text information based on user input and context. Using pre-designed templates, fill in the appropriate place with user input information to generate image-text content. According to the input keywords and patterns, generate appropriate text responses, commonly used in automatic SMS reply scenarios. Identify keywords in the input text, then generate related image-text content based on these keywords.
[0005] However, traditional intelligent auxiliary image-text generation technology has some defects:
[0006] (1) Lack of creativity and personalization: Traditional technology is usually based on predefined rules, templates or patterns, making it difficult to generate creative and personalized image-text content, limiting user experience.
[0007] (2) Limited flexibility: Traditional technology has limitations in adapting to diverse user needs and contexts, and cannot flexibly generate complex and variable image-text information.
[0008] (3) Lack of context understanding: Traditional techniques often struggle to deeply understand the context of user input, leading to generated text content that may lack accuracy and coherence.
[0009] To this end, a multi-home device interaction method for intelligent auxiliary generation of text and graphics and a storage medium are proposed. SUMMARY
[0010] Therefore, the embodiments of the present application aim to provide a multi-home device interaction method for intelligent auxiliary generation of text and graphics and a storage medium to solve or alleviate the technical problems existing in the prior art, i.e., lack of creativity and individuality, limited flexibility and lack of context understanding, and at least provide an advantageous choice for this;
[0011] The technical solution of the embodiments of the present application is as follows:
[0012] First aspect
[0013] A multi-home device interaction method for intelligent auxiliary generation of text and graphics: aims to realize a more intelligent, personalized and coherent home device interaction experience, which includes:
[0014] S1, prediction: using a cellular automaton model for interaction prediction:
[0015] At this stage, the present application uses a cellular automaton model to predict the interaction information that the user may output in the next time step. The cellular automaton model can predict the user's possible interaction based on the current state and user history interaction, based on predefined rules and conditions. The specific steps are as follows:
[0016] Cell state and input: Each cell represents the state of a device, and the user's historical interaction is encoded as input.
[0017] Rules and cell transition function: define a set of rules representing the transition rules of device state. Use the cell transition function to calculate the next state.
[0018] S2, update: using LSTM to generate interaction text:
[0019] At this stage, the present application uses LSTM to generate text information related to the predicted interaction to capture context information, and then outputs the corresponding text. The specific steps are as follows:
[0020] Input and hidden state: use the interaction predicted by the cellular automaton as input, combined with the previously generated text to form the output. Use the hidden state to capture context information.
[0021] Output: using the output layer, through weights and biases, and the Softmax function, calculate the output as the generated interaction text;
[0022] S3, Integration: Fusion of Cellular Automaton Prediction and LSTM Generation:
[0023] In this phase, the invention fuses the prediction results of the cellular automaton model with the text generated by LSTM to produce the final interrogative interactive text. The specific steps are as follows:
[0024] Fusion weight calculation: According to the prediction confidence and the generation confidence, the fusion weight is calculated to balance the influence of both;
[0025] Fusion result generation: Through linear weighting, the fusion result is generated according to the fusion weight as the final interrogative interactive text;
[0026] Through this method, the invention can combine the prediction of cellular automaton and the generation of LSTM to produce more coherent and personalized home device interactive text, while fully considering the context information and user needs, improving user experience. Compared with traditional technology, this method can better cope with complex situations and natural language expressions, providing higher quality image generation.
[0027] Second aspect
[0028] A storage medium, the storage medium stores program instructions for executing the intelligent auxiliary image generation of multiple home device interaction method as described above.
[0029] This storage medium can be a computer hard disk, solid state disk, memory, cloud storage, etc. The storage medium stores program instructions related to the intelligent auxiliary image generation of multiple home device interaction method, which describes the implementation details of the entire method, including the prediction, update and integration of each stage.
[0030] Specifically, the program instructions in the storage medium can include the following:
[0031] (1) Cellular automaton model implementation instructions: This part of the instructions covers the construction of the cellular automaton model, the definition of the state transition function, the setting of the rules, etc. It describes how to predict the user's possible interactive information in the next time step based on the current state and user history interaction.
[0032] (2) LSTM model implementation instructions: This part of the instructions describes the architecture of the LSTM model, the processing of the input and hidden state, the definition of the output layer, etc. It explains how to use LSTM to generate text information related to the predicted interaction.
[0033] (3) Fusion and Integration Instructions: These instructions cover how to calculate fusion weights, perform linear weighted fusion, and generate the final conversational interactive text. It describes how to fuse the prediction results of the cellular automaton model with the output of the LSTM model to generate optimized graphic content.
[0034] (4) Data Processing and Preprocessing Instructions: In practical applications, data processing and preprocessing are very important steps. These instructions describe how to prepare training data, encode input, perform feature engineering, etc.
[0035] (5) User Interface and Interaction Instructions: If user interface and interaction are involved, the instructions in the storage medium may also include how to receive user input, display generated interactive text, etc.
[0036] (6) These program instructions in the storage medium form the basis for implementing the entire intelligent auxiliary graphic generation multi-home device interaction method. By executing these instructions, the system can predict user interaction, generate interactive text, and fuse the outputs of the cellular automaton model and the LSTM model in practical applications, providing a more intelligent, personalized, and smooth home device interaction experience.
[0037] In summary, compared with the prior art, the beneficial effects of the intelligent auxiliary graphic generation multi-home device interaction method and storage medium provided by the present application are as follows:
[0038] I. Personalized Interaction: Through the prediction of the cellular automaton model, the system can better understand the user's historical interaction and habits, thereby generating more personalized and individualized interactive content that meets the user's preferences and individual needs. This personalized interaction can enhance the emotional connection between the user and the home device, improving user satisfaction. By providing a more personalized and coherent interactive experience, users are more likely to participate in the control and management of home devices, thereby increasing their positive engagement and willingness to use the devices. The cellular automaton model can generate certain creative content based on rules and conditions, which can surprise and entertain users. Compared with traditional methods, this method is more likely to generate novel and creative interactive content.
[0039] II. Coherence and Smoothness: Through the application of the LSTM model, the system can capture contextual information and generate more coherent interactive text. This helps ensure that the interactive content is more coherent and smooth in terms of semantics, improving the communication effect between the user and the device. The LSTM model has advantages in generating text and can better simulate the characteristics of human natural language expression. This means that the generated interactive text is more natural and easier for users to understand and accept, thereby reducing communication barriers.
[0040] III. Context awareness: The fusion process in the integration stage fully considers the prediction of the cellular automaton model and the generation of the LSTM model, enabling the system to better understand and integrate the current context when generating interactive text, improving the relevance and practicality of the interactive content. This method can flexibly adapt to different home devices and user needs, and by adjusting the rules and model parameters, it can achieve a wider range of applications. This allows the system to provide consistent and effective interactive support in different home environments. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.
[0042] Figure 1 Logical diagram of the present application;
[0043] Figure 2 Program flowchart of embodiment five of the present application;
[0044] Figure 3 Program flowchart of embodiment five of the present application;
[0045] Figure 4 Program flowchart of embodiment five of the present application; DETAILED DESCRIPTION
[0046] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below. In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the scope of the present application, therefore the present application is not limited to the specific embodiments disclosed below;
[0047] Please refer to Figure 1 , the specific embodiments will provide related technical solutions:
[0048] An intelligent auxiliary image-text generation multi-home device interaction method, comprising the following steps:
[0049] S1, prediction: using a cellular automaton model for interaction prediction:
[0050] In this stage, the present embodiment uses a cellular automaton model to predict the interaction information that the user is likely to output in the next time step. This is computed by applying cellular automaton rules based on the current state and information about the user's historical interactions. Specifically:
[0051] S1.1, Cell state and input:
[0052] Each device corresponds to a cell, which represents the current state of the device.
[0053] The user's historical interactions are encoded into an input vector, which is used to influence the prediction.
[0054] S1.2, Rules and cell transition function:
[0055] The present embodiment defines a set of rules R, which describe the transition behavior of the device state.
[0056] The cell transition function f computes the desired state in the next time step based on the rules and the current state, input.
[0057] Logical principle example: Suppose the present embodiment has a smart light, whose state can be on or off. The present embodiment wants to predict the user's likely operation in the next time step (e.g., turn on the light). The present embodiment can define a rule that if the current state is off and there is a command to turn on the light in the user's history, then the next state is predicted to be on.
[0058] S2, Update: Generate interaction text using LSTM
[0059] In this stage, the present embodiment uses an LSTM model to generate text information related to the predicted interaction, in order to capture contextual information, and then outputs the corresponding text:
[0060] S2.1, Input and hidden state:
[0061] The interaction predicted using the cellular automaton model is used as input, combined with the previously generated text.
[0062] The input is converted into a continuous vector through word embedding, and then input into the LSTM model.
[0063] S2.2, LSTM model:
[0064] The LSTM model updates the hidden state to capture the semantic information of the previous text.
[0065] The hidden state is updated at each time step, reflecting the evolution of the contextual information.
[0066] S2.3, Output:
[0067] The output layer is used to map the hidden state of the LSTM to a probability distribution over the vocabulary.
[0068] The Softmax function is used to convert the output into probabilities.
[0069] Logical principle example: Suppose this embodiment wants to generate a sentence to answer the user's question about the status of the lamp. In the input stage, this embodiment combines the word embeddings of the user's question and the word embeddings of the previously generated content into the LSTM. The LSTM will generate the corresponding answer text according to the context information and the semantics of the user's question.
[0070] S3, Integration: Fusion of Cellular Automaton Prediction and LSTM Generation
[0071] In this stage, this embodiment fuses the prediction results of the cellular automaton model and the output of the LSTM model to produce the final interactive text:
[0072] S3.1, Fusion weight calculation:
[0073] Using the confidence of the prediction and the confidence of the generation, the fusion weight Wt is calculated.
[0074] This weight represents the relative weight of using cellular automaton prediction and LSTM generation in the integration.
[0075] S3.2, Fusion result generation:
[0076] By linear weighting, the cellular automaton prediction result and the text generated by the LSTM are integrated into the final interactive text.
[0077] Logical principle example: Suppose this embodiment's cellular automaton predicts that the user may want to turn on the light. And the text generated by the LSTM is an encouraging sentence, reminding the user that they can adjust the brightness of the light as needed. Through the fusion weight, this embodiment can decide the final generated interactive text according to the importance of the cellular automaton prediction and the encouraging content of the LSTM generation, to provide a more heuristic reply.
[0078] In summary, the entire process includes cellular automaton prediction, LSTM generation and fusion, in order to produce more personalized, coherent and creative graphic content in multi-home device interaction.
[0079] Further, the generated interactive text:
[0080] After obtaining the data text of the interactive text, the method of generating the graphic text can be selected:
[0081] (1) Using a text-to-image library: Use some ready-made text-to-image libraries that provide a simple API interface. By inputting the text into the API, you can output the picture, and you can also convert the input text into a specified image format (PNG or JPEG). Preferred text-to-image libraries include ImageMagick, Pillow (Python), and GraphicsMagick, etc.
[0082] (2) Command-line tools: Command-line tools can also be used to convert text to pictures. For example, use the command-line tools pandoc and ImageMagick on Linux or Mac operating systems to directly convert text files into interactive pictures.
[0083] The technical features of the above specific embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above specific embodiments are described, but as long as the combination of the technical features does not exist Contradiction, it should be considered as the scope of the present disclosure.
[0084] Embodiment one
[0085] According to the above specific embodiments and examples, the present embodiment further provides the following technical solutions:
[0086] The cellular automaton model predicts the next possible interaction according to the current state and user history interaction, and the cellular automaton model:
[0087] State Cell: Ct represents the state at time step t.
[0088] Input: Xt represents the input at time step t, including user history and current context; Ut represents the user interaction information at time step t;
[0089] Rules: R represents the rule set of the cellular automaton, which represents the conversion rule of the device state; For example, the rule can specify that if the user mentions "turn on the light", the state becomes "on".
[0090] Next state Next Cell: represents the state of Cell at time step t+1, which is calculated according to the current state, input and rule; The core logic of the cellular automaton model is to predict the next state given the current state, input and rule. This can be represented by the following formula:
[0091]
[0092] f is a function of state transition rule, which defines how to calculate the next state according to the current state, input and rule. Specifically, f can be a logical function, a mapping function, a rule set, etc., and the specific form depends on the cellular automaton model you design.
[0093] Cell State: Each cell represents the state of a device, which can be the brightness, color, etc. of a lamp, Si_t represents the state of cell i at time step t;
[0094] Cell Transition Function: Map the current state, input and rule to the next state:
[0095]
[0096] wherein Transition is a defined cell transition function that calculates the next state according to the current state, user interaction information and rule, and the next state is equal to the desired state. The Transition function can be implemented based on a rule base, a machine learning model or other methods. For example, a rule base can be used to match user interaction and perform corresponding state transition. Or a machine learning model can be used to predict state transition, such as using a classifier to predict the next state according to user interaction.
[0097] Preferably, the logical function f: it determines the next state according to the current state and input. Suppose this embodiment has a smart door lock device, the state Ct represents the locking state of the door (1 for locking, 0 for unlocking), the user history and current context are encoded into the input Xt, and the user's interaction information at time step t is Ut. This embodiment can define the logical function f as follows:
[0098]
[0099]
[0100]
[0101]
[0102] Preferably, the mapping function f: it maps the value of the next state according to the current state and input. Suppose this embodiment has a smart window curtain device, the state Ct represents the opening degree of the curtain (0 for closed, 1 for fully open), the user history and current context are encoded into the input Xt, and the user's interaction information at time step t is Ut. This embodiment can define the mapping function f as follows:
[0103]
[0104]
[0105]
[0106]
[0107] Preferably, the state transition function f: a set of rules that computes the next state based on a set of rules. Suppose this embodiment has a smart speaker device, the state Ct represents the volume level (0-10), the user history and current context are encoded into the input Xt, the user's interaction information at time step t is Ut. This embodiment can define the rule set R as follows:
[0108]
[0109]
[0110]
[0111] Preferably, the Transition function is implemented based on a rule base, the training of the rule base is performed as follows:
[0112] P1, Rule base definition: First define a rule base, which contains different rules. Each rule is associated with a state transition case, such as state transition according to the user's instruction.
[0113] P2, User interaction matching: At each time step, the Transition function matches the user's interaction information with the rules in the rule base. If the user's interaction information meets the conditions of a certain rule, the rule will be triggered.
[0114] P3, State transition: The triggered rule will specify the way of state transition. According to the matched rule, the Transition function will calculate the state change defined in the rule to determine the next state.
[0115] P4, Update state: Finally, the Transition function applies the calculated next state to the cellular automaton model, updating the state of the device, thus realizing the state transition.
[0116] Exemplary Transition function:
[0117] Rule base example:
[0118] If the user mentions "turn on the light", the state becomes "on".
[0119] If the user mentions "turn off the light", the state becomes "off".
[0120] def Transition(current_state, user_interaction, rule_library):
[0121] next_state = current_state # Default state remains unchanged
[0122] for rule in rule_library:
[0123] if rule.condition(user_interaction): # Check if the rule's condition is met.
[0124] next_state = rule.action(current_state) # Perform a state transition according to the rule.
[0125] return next_state
[0126] The example above demonstrates a rule base where the `Transition` function iterates through each rule, checking if the user interaction information meets the rule's conditions. If it does, it executes the corresponding state transition operation to obtain the next state. Through this rule base-based implementation, the `Transition` function can flexibly perform state transitions based on different user interaction scenarios, thus realizing the state prediction and transition functions in the cellular automata model.
[0127] Example 2
[0128] Based on the above specific implementation methods and embodiments, this embodiment further provides the following technical solutions:
[0129] Use an LSTM-based text generation model to generate interactive query text, including:
[0130] Input: It represents the input at time step t, including the interaction predicted by the cellular automaton and the previously generated text;
[0131] Hidden state: Ht represents the hidden state at time step t, used to capture context information.
[0132] Output: Ot represents the output at time step t, as the generated interactive text.
[0133] In S2, the input: It represents the input at time step t, including the interaction predicted by the cellular automaton and the previously generated text.
[0134]
[0135] where E denotes the word embedding function, Ut represents the predicted interaction by the cellular automaton, and Gt represents the generated text up to time step t. At time step t, the hidden state is used to capture the contextual information. It is obtained by updating the LSTM model with the previous hidden state and the input, which includes the predicted interaction by the cellular automaton and the previously generated text. At time step t, the output is obtained by performing a linear transformation on the hidden state and applying the Softmax operation, representing the probability distribution of the generated interaction text.
[0136] Hidden state: Ht represents the hidden state at time step t, which is used to capture the contextual information:
[0137]
[0138] where LSTM denotes the updating process of the LSTM model, and Ht-1 represents the hidden state at time step t-1.
[0139] Output: Ot represents the output at time step t, which is the generated interaction text:
[0140]
[0141] where Wo and bo are the weights and bias of the output layer, and Softmax represents the Softmax function, which converts the linearly transformed results into a probability distribution.
[0142] Through this process, we can generate the interaction text at the next time step using the LSTM model based on the previous generated content and the predicted interaction by the cellular automaton. The entire process includes the representation of the input, the updating of the hidden state, and the calculation of the output, to achieve the generation of the inquiry-based interaction text based on LSTM.
[0143] Specifically, when using the Softmax function: a set of real values is converted into a probability distribution. Given a real number vector z = [z_1, z_2,..., z_n], the Softmax function converts each element zi into a probability value p_i, such that the sum of all probability values is equal to 1. The formula of the Softmax function is as follows:
[0144]
[0145] where e is a natural constant (approximately equal to 2.71828). The Softmax function exponentiates each element in the input vector and then divides the exponentiated values by the sum of the exponents of all elements. This ensures that each probability value is between 0 and 1, and the sum of all probability values is 1, forming a probability distribution. Based on the properties of the exponential function, its main goal is to standardize the input real values so that they can be interpreted as probabilities:
[0146] P1, Amplify differences: The Softmax function exponentiates each element of the input vector, which amplifies larger values while compressing smaller ones, increasing the differences between them.
[0147] P2, Standardize probabilities: The sum of exponents in the denominator will ensure that the sum of the output probability values is 1. This means that the Softmax function can convert the original real values into a relative probability distribution, so that they can be used to represent the probabilities of different categories.
[0148] P3, Class selection: The Softmax function makes the class corresponding to the largest probability value in the output the most likely class, as it responds more strongly to larger input values.
[0149] P4, The logic of the Softmax function makes it well-suited for classification tasks, as it effectively maps the original input to a probability distribution, helping the model make predictions and decisions. In the LSTM-based text generation model, the Softmax function is used to convert the linear transformation results of the LSTM model output into a probability distribution of the generated interactive text. In this way, we can select the text with the highest probability as the final generation result.
[0150] Example Three
[0151] According to the specific embodiments and examples described above, this embodiment further provides the following technical solutions:
[0152] Use word embeddings to map the input discrete symbols, interactive and text to continuous vector representations, then update the hidden state through the LSTM model, and finally get the generated interactive text through the output layer.
[0153] Specifically, use word embeddings to map discrete symbols:
[0154] P1, Input representation: Our model receives the interactive and previously generated text predicted by the cellular automaton as input, which are usually discrete symbols (such as words). In order to process them in the model, we first need to map these discrete symbols to continuous vector representations.
[0155] P2, Word Embedding: We use word embedding techniques such as Word2Vec, GloVe, or BERT to map discrete symbols into continuous vector space. In this way, each discrete symbol is represented as a real-valued vector that preserves the semantic relationships between symbols. This allows the model to better capture the semantic information of the text.
[0156] P3, LSTM Model Hidden State Update:
[0157] Input Combination: After word embedding, we get the continuous vector representation of the interaction predicted by the cellular automaton and the previously generated text. These vectors will be used as input to the LSTM model.
[0158] LSTM Model: LSTM (Long Short-Term Memory) is a type of recurrent neural network that is particularly suitable for processing sequential data. It has a hidden state that can capture the context information in the sequence.
[0159] Hidden State Update: At each time step, the LSTM model receives the input at the current time step and the hidden state at the previous time step, and then calculates the new hidden state. This hidden state will contain the context information from the previous time steps of the input sequence, helping the model understand the context of the text.
[0160] P4, Output Layer Generates Interaction Text:
[0161] LSTM Model Output: At each time step, the hidden state of the LSTM model will be used as output to capture the context information at the current time step.
[0162] Linear Transformation: We pass the output of the LSTM model through a linear transformation (weighted sum) to the output layer. This linear transformation maps the model output to a representation more suitable for generating interactive text.
[0163] Softmax Probability Distribution: After linear transformation, we use the Softmax function to convert the model output into a probability distribution. Each possible interactive text will be associated with a probability value indicating the likelihood of generating that text.
[0164] Through this process, we can map discrete interactions and text inputs into continuous vector representations, capture context information and update states through the LSTM model, and finally get the probability distribution of the generated interactive text through the output layer. This allows us to select the interactive text with the highest probability as the final generated result based on the model's prediction.
[0165] Embodiment Four
[0166] According to the specific embodiments and examples described above, this embodiment further provides the following technical solutions:
[0167] In S3 of this embodiment: Linearly weighted fusion is used: At this stage, we will combine the prediction results of the cellular automaton model and the output of the LSTM model through linear weighting to generate the final interrogative interactive text. Here we involve two key aspects: fusion weight and fusion result.
[0168] Specifically, the goal of fusion is to obtain more accurate, interesting and appropriate home device interactive text by integrating the outputs of two different models. The cellular automaton model may predict interactive information more based on rules and historical data, while the LSTM model can better capture the context and generate natural language.
[0169] Fusion weight: At time step t, we use fusion weight Wt to balance the interactive information predicted by the cellular automaton and the text information generated by the LSTM. The fusion weight is calculated by the following formula:
[0170]
[0171] Where α is a parameter that controls the weight balance, Confidence represents the confidence assessment of prediction or generation; the larger α is, the more inclined to use the interactive information predicted by the cellular automaton, the smaller α is, the more inclined to use the text information generated by the LSTM;
[0172] Fusion result: using the fusion weight, the prediction results of the cellular automaton model and the output of the LSTM model are combined through linear weighting:
[0173]
[0174] The fusion result Ft is the final interrogative interactive text generated at time step t. It integrates the information of two models and selectively fuses the cellular automaton prediction and LSTM generated text according to the weight balance;
[0175] Ot is the interactive text output generated by the LSTM, and Ct+1 is the next state predicted by the cellular automaton. The fusion result Ft is the final interrogative interactive text generated at time step t. Through the adjustment of the fusion weight, we can balance the contribution between the cellular automaton model and the LSTM model to a certain extent, so as to produce more comprehensive and adaptive interactive text. Through this fusion process, we can effectively integrate the information from the cellular automaton and the LSTM model to generate more accurate, interesting and appropriate home device interactive text. The logic of this process helps to improve the intelligence and user experience of the system in practical applications.
[0176] Further, the logic principle of the fusion process is to integrate the advantages of different models, control the weight distribution of both using fusion weights, and achieve a better balance when generating interactive text. This enables the model to selectively fuse the outputs of different models according to task requirements and the quality of generated results, thus generating more creative, accurate, and personalized home device interactive text. This fusion method has potential in improving system performance and user experience.
[0177] Further, regarding α: initially, you can set the initial value of α based on experience or intuition. For example, if you think that the cellular automaton prediction is more reliable, you can set a larger initial value; if you think that the LSTM generation is more accurate, you can set a smaller initial value. Then, adjust according to the actual effect. During the training and validation process, you can use cross-validation to determine the optimal α value. Divide the training data into training set and validation set, try different α values and evaluate the quality of generated results on the validation set, select the α value that performs best on the validation set. You can consider using adaptive methods to dynamically adjust the value of α. For example, according to the model's training progress, the confidence of generated text, and other factors, automatically adjust the size of α. This may require some additional work, but can adapt to the needs of the model at different stages. Use hyperparameter optimization techniques such as Bayesian optimization or grid search to search for the optimal α value. These methods can automatically search for the best parameter configuration within a certain range to achieve better performance. If you have multiple cellular automaton models and LSTM models, you can consider using model ensemble methods. In this case, you can set different α values for each model, then fuse their outputs by weighted average or voting, etc. If user feedback is available, you can adjust the α value according to user preferences and satisfaction. For example, users may prefer the generated interactive text, or trust the cellular automaton prediction more.
[0178] Embodiment Five
[0179] According to the specific embodiments and examples described above, this embodiment further provides the following technical solutions:
[0180] This embodiment further provides a storage medium, which stores a multi-home device interaction method for implementing intelligent auxiliary image-text generation as described in embodiments one to four above. Please refer to Figures 2-4 which is in the form of C++ pseudo code, and the principle is as follows:
[0181] S1: Cellular automaton prediction:
[0182] This section demonstrates how to use the cellular automaton model to predict the device state at the next time step. Cellular automata is a mathematical model that evolves in discrete time and space, used to describe the evolution of states. Here, we predict the next state by calling the CellAutomatonPredict function based on the current state, user input, and rules. This embodies the concept of state prediction through rules and historical data.
[0183] S2: LSTM Model Update:
[0184] This section shows how to use an LSTM-based model to update the hidden state and generate text. LSTM is a type of recurrent neural network used to process sequential data. In the code, the LSTMUpdate function represents how to update the hidden state based on the input and the previous hidden state. This process captures contextual information, enabling better understanding of the context and generating appropriate text.
[0185] S3: Linear Weighted Fusion:
[0186] This section shows how to combine the cellular automaton predicted state and the LSTM generated text through linear weighted fusion to generate the final interactive text. The LinearWeightedFusion function calculates the fusion result based on the predicted state, LSTM output, and fusion weights. The fusion weight is controlled by the parameter α, which can be adjusted according to model performance and requirements to balance the contribution of both types of information.
[0187] This code shows how to combine a cellular automaton model and an LSTM-based text generation model to generate text suitable for home device interaction through an example. Each stage represents a key step in the entire process, covering prediction, context capture, and information fusion to achieve a more intelligent, interesting, and personalized home device interaction experience. This example pseudocode can serve as a starting point for actual implementation, customized and extended according to specific requirements.
Claims
1. A method for intelligent auxiliary graphic text generation and multi-home device interaction, characterized in that: Comprising the following steps: S1, prediction: using a cellular automaton model, predicting the possible interaction information of the user in the next time step according to the current state and the user's historical interaction, calculating the expected state according to the current state, input and rules, and outputting the expected state; S2, update: using LSTM, inputting the expected state in S1, including the interaction predicted by the cellular automaton and the previously generated text, and outputting the text after capturing the context information; S3, integration: integrating the prediction result of the cellular automaton model and the output of the LSTM text generation model in a linear weighted form to generate the final inquiry interaction text; In S1, the cellular automaton model predicts the next possible interaction according to the current state and the user's historical interaction, and the cellular automaton model: State Cell: Ct represents the state at time step t; Input: Xt represents the input at time step t, including the user's history and the current context; Rules: R represents the rule set of the cellular automaton, which represents the conversion rule of the device state; Next Cell: represents the state of Cell at time step t+1, which is calculated according to the current state, input and rules: ; f is the function of state conversion rule, which defines how to calculate the next state according to the current state, input and rules; In S2, input: It represents the input at time step t, including the interaction predicted by the cellular automaton and the previously generated text: ; Where E represents the word embedding function, Ut represents the interaction predicted by the cellular automaton, and Gt represents the text generated before time step t; Hidden state: Ht represents the hidden state at time step t, which is used to capture context information: ; Where LSTM represents the update process of the LSTM model, and Ht-1 represents the hidden state at time step t-1; Output: Ot represents the output at time step t, which is the generated interaction text: ; Where Wo and bo are the weights and biases of the output layer, and Softmax represents the Softmax function. 2.The multi-home device interaction method of claim 1, wherein: In S2: Cell State: Each cell represents the state of a device, and Si_t represents the state of cell i at time step t; Cell transition function: maps the current state, input and rules to the next state: ; Ut represents the user interaction information at time step t; Where Transition is the defined cell transition function, which calculates the next state according to the current state, user interaction information and rules, and the next state is equal to the expected state. 3.The multi-home device interaction method of claim 1, wherein: In S2: an LSTM-based text generation model is used to generate inquiry interaction text, including: Input: It represents the input at time step t, including the interaction predicted by the cellular automaton and the previously generated text; Hidden state: Ht represents the hidden state at time step t, which is used to capture context information; Output: Ot represents the output at time step t, which is the generated interaction text. 4.The multi-home device interaction method of claim 3, wherein: The input discrete symbols, the interaction and the text are mapped to continuous vector representations using word embeddings, then the hidden state is updated by an LSTM model, and finally the generated interactive text is obtained by an output layer. 5.The multi-home device interaction method of claim 2, wherein: In the S3, linear weighted fusion is used: The fusion weight: Wt represents the fusion weight at time step t: ; Wherein, α is a parameter for controlling the balance of weights, Confidence represents the confidence evaluation of prediction or generation; the larger α is, the more inclined to use the interactive information predicted by the cellular automaton, the smaller α is, the more inclined to use the text information generated by the LSTM; The fusion result: ; Wherein, Ct+1 is the next state predicted by the cellular automaton; the fusion result Ft is the final interrogative interactive text generated at time step t.
6. A computer-readable storage medium, characterized in that: The computer readable storage medium stores program instructions for executing the multi-home device interaction method according to any one of claims 1-5.
Citation Information
Patent Citations
Method, system and device for determining user scene and storage medium
CN111857331A
Natural language understanding method and device fusing dialogue context information
CN116542256A