Smart home multi-modal dialogue method, device, equipment and medium

By extracting word features and context vectors through a multimodal dialogue model and combining them with a knowledge base, the problem of mismatch between logic and emotional expression in multimodal dialogue technology is solved, generating intelligent responses that conform to grammar and emotion, thus improving the human-likeness and naturalness of the dialogue experience.

CN120822527BActive Publication Date: 2025-12-26QINGDAO TAPER ROBOTICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511331824.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-26
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing multimodal dialogue technologies struggle to meet human emotional needs, generating responses that are not logically consistent in content and do not match the user's emotional state in terms of emotional expression, and lack effective modeling of explicit and implicit knowledge.

Method used

By acquiring device dialogue data, text enhancement is performed by extracting word features, part-of-speech features, and relational features of adjacent words using a multimodal dialogue model. Combined with context vectors and a knowledge base, grammatically and logically consistent responses are generated, and user emotional states are predicted to generate more empathetic and adaptive replies.

Benefits of technology

The generated responses are logical in content and the emotional expression matches the user's state, enhancing the anthropomorphism of multimodal dialogue technology and the naturalness of the dialogue experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822527B_ABST
    Figure CN120822527B_ABST
Patent Text Reader

Abstract

The present application relates to the field of household appliances, and provides a smart home multi-modal dialogue method, device, equipment and medium, the method comprises: inputting device dialogue data into a multi-modal dialogue model, extracting word features of text information of the device dialogue data, part-of-speech features corresponding to the word features and relationship features between adjacent word features, respectively performing text enhancement on the word features and the part-of-speech features according to the relationship features, extracting context vectors according to the text information, obtaining a decoding word sequence in combination with the part-of-speech features and the word features after text enhancement, predicting a user emotional state according to the device dialogue data, and generating a reply prediction result in combination with the decoding word sequence. The present application solves the defect that multi-modal dialogue technology is difficult to match human emotional needs, and can ensure that the generated reply not only meets the logic in content, but also matches the user state in emotional expression.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of household appliances, and in particular to a smart home multi-modal dialogue method, device, equipment and medium. BACKGROUND

[0002] Multi-modal dialogue technology has been widely used in smart home, e-commerce customer service, enterprise marketing and other complex scenarios. The development of natural language processing and deep learning technology has promoted the progress of multi-modal dialogue technology, and has become a new trend of human-computer interaction.

[0003] Current multi-modal dialogue technology focuses on enhancing multi-modal context understanding by using collected knowledge, but lacks exploration in knowledge-guided text generation to improve the humanization of multi-modal dialogue technology. The lack of human emotional knowledge makes it difficult to analyze and track the emotional changes of users in the dialogue, resulting in generated replies that are difficult to meet human emotional needs. In the aspect of proper emotional intelligence expression, it is in the blank or initial stage, making the generated content present "mechanized" characteristics. At the same time, explicit knowledge and implicit knowledge are different in characteristics and effects, which restricts the extraction and mining ability of multi-modal dialogue technology for knowledge. How to model according to the characteristics of the two types of knowledge and enhance the text personification from the aspects of grammar, logic and emotion is a key problem to be solved. SUMMARY

[0004] The present application provides a smart home multi-modal dialogue method, device, equipment and medium to solve the defect that the multi-modal dialogue technology in the prior art is difficult to meet the human emotional needs, and to ensure that the generated reply not only meets the logic in content, but also matches the user state in emotional expression.

[0005] The present application provides a smart home multi-modal dialogue method, comprising: obtaining device dialogue data; inputting the device dialogue data into a multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model; wherein the multi-modal dialogue model is trained according to device historical dialogue data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical dialogue data; the multi-modal dialogue model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to the text information in the input device dialogue data, and to perform text enhancement on the word features and part-of-speech features according to the relationship features, respectively, to extract context vectors according to the text information, and to decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, to predict a user emotional state according to the device dialogue data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0006] It should be noted that by inputting the obtained device dialogue data into the multi-modal dialogue model, the relationship features between the word features, the part-of-speech features and the adjacent word features are extracted, so as to perform text enhancement on the word features and the part-of-speech features based on the relationship features, enhance the explicit grammatical constraints, enable the model to more deeply understand the semantics, grammar and context association of the input text, make the generated text more fluent, ensure that the generated reply is associated with the entire dialogue history through the extraction of the context vector, and decode the part-of-speech and word features after text enhancement, which is helpful to generate a grammatically correct and logically coherent reply framework, and integrate knowledge base information to improve the accuracy and knowledge of the reply, and gradually generate words in the decoding process to form a complete reply sequence, then predict the user emotional state according to the text and image information, identify the emotional state of the user, so as to generate a reply with more empathy and adaptability, so as to combine language logic (decoded word sequence) and emotional information (emotional state), and ensure that the generated reply not only meets the logic in content, but also matches the user state in emotional expression.

[0007] According to the present application, a smart home multi-modal dialogue method and a multi-modal dialogue model are provided, which comprise: a text enhancement layer for extracting word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features from text information in input device dialogue data, performing text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature according to a previous word feature, a part-of-speech feature corresponding to the previous word feature and relationship features between the previous word feature and the subsequent word feature, determining corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtaining text-enhanced word sequences and text-enhanced part-of-speech sequences; a logic decoding layer for extracting a context vector from the text information in the device dialogue data, and decoding the text-enhanced part-of-speech sequences and the text-enhanced word sequences to obtain a decoded word sequence; an emotional expression layer for predicting a user emotional state according to the text information and image information in the device dialogue data to obtain an emotional prediction state; and a reply generation layer for fusing the decoded word sequence and the emotional prediction state, and predicting a reply based on the fused features to obtain a reply prediction result.

[0008] It should be noted that the input text information is extracted by using the text enhancement layer, the word features, the part-of-speech features and the relationship features are introduced, the understanding of the input text by the model is more in-depth and accurate, the current word is enhanced by using the word features and the relationship features of the previous word, the explicit language rules such as part-of-speech are used, the language knowledge constraint is enhanced, the grammatical meaning and the fluency of the language expression of the text enhancement word sequence and the text enhancement part-of-speech sequence are enhanced, the generated text is more fluent, the text enhancement word sequence and the text enhancement part-of-speech sequence are used explicitly, the model generates the reply while following the grammatical rules, and the context vector extracted based on the text information is combined, so that the generated reply is consistent with the dialogue history in semantics, and the content generated conforms to the grammar and the logic, thereby the explicit knowledge guidance is strengthened, the logical meaning of the generated decoding word sequence and the focusing degree of the dialogue context are enhanced, and the emotional clues are extracted from the text and the image information, so that the key implicit information of the user emotion is accurately captured, the implicit knowledge guidance is enhanced, the text semantic and the emotion related weight are dynamically balanced, the generated text has rich emotional expression, the emotional meaning of the generated text and the human touch in the dialogue process are enhanced, so that the emotional implicit knowledge and the logical content are used to help the model generate a reply with more empathy, more appropriate and more resonant, so that the dialogue experience is more natural and more humanized.

[0009] According to the intelligent home multi-modal dialogue method provided by the application, the context vector is extracted according to the text information in the device dialogue data, and the text enhancement part-of-speech sequence and the text enhancement word sequence are combined for decoding to obtain a decoding word sequence, comprising: extracting the context vector of each text-enhanced word feature corresponding to the text enhancement word sequence according to the text information in the device dialogue data; for each text-enhanced word feature, determining the hidden state corresponding to the text-enhanced word feature according to the context vector of each text-enhanced word feature, the text enhancement part-of-speech sequence and the text enhancement word sequence, combining the knowledge base, and predicting the probability distribution of the next word to obtain the decoding word sequence; wherein the knowledge base is constructed based on the query vector of each word and the knowledge entity of each word.

[0010] It should be noted that by extracting the context vector corresponding to each word feature, the context information is captured to provide a vector representation reflecting the position of each word in the entire dialogue or sentence, which helps the decoder to maintain semantic and logical coherence when generating subsequent words, and the context vector of the word, the text enhancement part-of-speech feature and the text-enhanced word feature of the word itself are used as input, and the pre-constructed knowledge base is introduced to further constrain the consistency of the reply content and the known facts, thereby determining the corresponding hidden state to guide the generation of the next word, which helps the model to generate a reply that conforms to the grammatical rules and the logical order.

[0011] According to the smart home multi-modal dialogue method provided by the application, the context vector of the enhanced word feature of each text, the text enhanced part-of-speech sequence and the text enhanced word sequence are combined with the knowledge base to determine the hidden state corresponding to the enhanced word feature of the text, and the probability distribution of the next word is predicted to obtain a decoding word sequence, including: SA1, when the enhanced word feature of the text is the first word in the text enhanced word sequence, the corresponding context vector is taken as a query vector to search the knowledge base to obtain a corresponding knowledge vector, and the knowledge vector is taken as the corresponding hidden state, and the probability distribution of the next word is predicted by using the hidden state; SA2, when the enhanced word feature of the text is not the first word in the text enhanced word sequence, the hidden state of the current word is determined according to the hidden state corresponding to the previous word, the enhanced word feature of the text, the enhanced part-of-speech feature of the text, the context vector and the knowledge vector, and the probability distribution of the next word is predicted in combination with the probability distribution of the current word predicted in the previous step, and the hidden state corresponding to the previous word is taken as the query vector of the enhanced word feature of the text to search the knowledge base to obtain the knowledge vector corresponding to the enhanced word feature of the text; SA3, the step SA2 is repeatedly executed until the probability distribution of the last word in the text enhanced part-of-speech sequence is predicted to obtain a decoding word sequence.

[0012] It should be noted that, for the first word, the context vector thereof is directly used for knowledge base query to ensure that the knowledge retrieval is closely related to the overall dialogue background, and the query result is taken as the initial hidden state, so that the beginning of the reply can be strongly affected by relevant knowledge, which helps to generate a starting word that is highly relevant to the current dialogue context and rich in information; for other words, the hidden state of the current word is determined according to the hidden state corresponding to the previous word, the enhanced word feature of the text, the enhanced part-of-speech feature of the text, the context vector and the knowledge vector, so as to strengthen the dependency relationship between words by explicitly passing the hidden state of the previous word, which helps to capture long-distance semantic association, so that each generated word is more accurate and reasonable, and the probability distribution of the current word predicted in the previous step is combined to make the prediction of the current step more stable and reliable, the step SA2 is repeatedly executed until the probability distribution of the last word in the text enhanced part-of-speech sequence is predicted, so as to ensure that a complete reply sequence matching the length of the input sequence can be generated, thereby generating a preliminary reply sequence that conforms to the language logic and contains rich knowledge content, laying a solid foundation for subsequent emotional fusion and final reply generation.

[0013] According to the intelligent home multi-modal dialogue method provided by the application, before the device dialogue data is input into the multi-modal dialogue model, the device historical dialogue data and the answer label corresponding to the device historical dialogue data, the reference word sequence, and the word label, the part-of-speech label, and the emotion label corresponding to each word in the reference word sequence are obtained; wherein the device historical dialogue data includes historical text information, historical image information, historical video information, and historical audio information corresponding to a target historical time; according to the historical video information, the historical picture frame sequence is obtained, and the historical picture frame sequence and the historical image information are normalized to obtain picture training data; the historical audio information is text recognized to obtain corresponding historical text, and the historical text and the historical text information are normalized to obtain text training data; according to the picture training data and the text training data, the device dialogue training data is obtained; the device dialogue training data is used as input data for training, the word label, the part-of-speech label, the reference word sequence, the emotion label, and the answer label are used as labels for training, the model to be trained is trained, and the multi-modal dialogue model is obtained.

[0014] It should be noted that by obtaining the device historical dialogue data and its label, the continuous video stream is decomposed into discrete picture frames, and the historical image information is normalized to facilitate unified input, eliminate differences in brightness, contrast, and the like of different pictures, make the model more robust to different inputs, speed up the convergence speed, and convert the audio information into text, so that all inputs are unified into a mode that is easy for the model to process, facilitating subsequent multi-modal fusion, and the historical text is normalized to ensure the consistency of the text data, which helps the model to learn more stable language representation, so that the picture training data and the text training data of different modalities are integrated into the same multi-modal training sample pair, the model is trained based on real multi-modal dialogue data and corresponding comprehensive labels, the model can learn the complete mapping relationship from input to output, learn more rich and robust feature representation, improve the overall performance, transfer the knowledge learned in the past dialogue to future interaction, and improve the generalization ability and practicality of the model.

[0015] According to the smart home multi-modal dialogue method provided by the application, the multi-modal dialogue model comprises a text enhancement layer, a logic decoding layer, an emotional expression layer and a reply generation layer; device dialogue training data is used as input data for training, word labels, word type labels, reference word sequences, emotional labels and answer labels are used as labels for training, a model to be trained is trained to obtain a multi-modal dialogue model, comprising: inputting text training data in the device dialogue training data into the text enhancement layer to obtain text enhancement training word sequences and text enhancement training word type sequences output by the text enhancement layer; inputting the text training data, the text enhancement training word sequences and the text enhancement training word type sequences into the logic decoding layer to obtain a decoding training word sequence output by the logic decoding layer; inputting the device dialogue training data into the emotional expression layer to obtain an emotional training state output by the emotional expression layer; inputting the decoding training word sequence and the emotional training state into the reply generation layer to obtain a reply training result output by the reply generation layer; constructing a first loss function according to the text enhancement training word sequences, the text enhancement training word type sequences, the word labels and the word type labels; constructing a second loss function according to the decoding training word sequence and the reference word sequence; constructing a third loss function according to the emotional training state and the emotional label; constructing a fourth loss function according to the reply training result and the answer label; obtaining a total loss function according to the first loss function, the second loss function, the third loss function and the fourth loss function, and ending the training based on the convergence of the total loss function.

[0016] It should be noted that the text part in the device dialogue training data is input into the text enhancement layer of the model for language feature extraction and enhancement, the word features, word type features and relationship features are extracted and enhanced to generate a representation that is richer, more semantic and grammatical than the original text, and the original text training data, the text enhancement training word sequences and the text enhancement training word type sequences are input into the logic decoding layer to combine the context vector and the structured text enhancement features to make a prediction that is more consistent with the logic and grammar, making the decoding process more robust and accurate, the device dialogue training data is input into the emotional expression layer to combine the text content and the image content to comprehensively judge the emotional state of the user, which helps the model to learn to capture and express emotions and convert complex emotional states into internal representations that can be processed by the model, so that the language logic (decoding training word sequence) and emotional expression (emotional training state) jointly determine the final generated reply, making the reply more intelligent, more natural and more human, and through multiple loss functions, the performance of the model in each subtask is more comprehensively evaluated, and through joint optimization, the different parts of the model can coordinate with each other to improve the overall performance.

[0017] According to the intelligent home multi-modal dialogue method provided by the application, the device dialogue data is obtained, including: obtaining device dialogue data; wherein the device dialogue data includes at least one of audio information, video information, image information and text information; when the device dialogue data includes audio information, the audio information is subjected to text recognition to obtain text information, and the device dialogue data is updated by using the text information; when the device dialogue data includes video information, according to the video information, a picture frame sequence is obtained, and the device dialogue data is updated by using the picture frame sequence.

[0018] It should be noted that by obtaining the device dialogue data, various original information generated by the device interaction can be captured, any potential useful data source is ensured not to be missed, and the audio information and video information are converted into structured text and image forms, so that subsequent data analysis, understanding, retrieval and application become more efficient and possible.

[0019] The application further provides an intelligent home multi-modal dialogue device, including: a data acquisition module, which acquires device dialogue data; a reply prediction module, which inputs the device dialogue data into a multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model; wherein the multi-modal dialogue model is trained according to device historical dialogue data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical dialogue data; the multi-modal dialogue model is used for extracting word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to text information in the input device dialogue data, and performing text enhancement on the word features and the part-of-speech features according to the relationship features, extracting context vectors according to the text information, and decoding by combining the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, predicting a user emotion state according to the device dialogue data, and generating a reply prediction result according to the predicted user emotion state and the decoded word sequence.

[0020] It should be noted that the device dialogue data obtained by the data acquisition module is input into the multi-modal dialogue model by the reply prediction module to extract the relationship features between the word features, the part-of-speech features and the adjacent word features, so as to perform text enhancement on the word features and the part-of-speech features based on the relationship features, enhance the explicit grammatical constraints, enable the model to more deeply understand the semantics, grammar and context connection of the input text, make the generated text more fluent, and ensure that the generated reply is associated with the entire dialogue history through the extraction of the context vector, decoding in combination with the part-of-speech and word features after text enhancement, helping to generate a grammatically correct and logically coherent reply framework, and integrating knowledge base information to improve the accuracy and knowledge of the reply, and gradually generating words in the decoding process to form a complete reply sequence, and then predicting the user emotional state according to the text and image information, identifying the emotional state of the user, so as to generate a reply with more empathy and adaptability, so as to combine language logic (decoded word sequence) and emotional information (emotional state), and ensure that the generated reply not only meets the logic in content, but also matches the user state in emotional expression.

[0021] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the smart home multi-modal dialogue method according to any one of the above when executing the computer program.

[0022] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the steps of the smart home multi-modal dialogue method according to any one of the above. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0024] Figure 1 is a flowchart of the smart home multi-modal dialogue method provided by the application;

[0025] Figure 2 is a flowchart of obtaining a decoded word sequence provided by the application;

[0026] Figure 3 is a flowchart of training a multi-modal dialogue model provided by the application;

[0027] Figure 4 is a structural diagram of the smart home multi-modal dialogue device provided by the application;

[0028] Figure 5 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0031] The present application will be described below in conjunction with Figure 1 A flowchart of an intelligent home multi-modal dialogue method of the present application is described below. The method comprises the following steps.

[0032] S11, obtaining device dialogue data;

[0033] S12, inputting the device dialogue data into a multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model; wherein the multi-modal dialogue model is trained according to device historical dialogue data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical dialogue data; the multi-modal dialogue model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to text information in the input device dialogue data, and to perform text enhancement on the word features and the part-of-speech features respectively according to the relationship features, to extract context vectors according to the text information, and to decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, to predict a user emotional state according to the device dialogue data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0034] It should be noted that the step numbers "S11-S12" in the present specification do not represent the sequence of the intelligent home multi-modal dialogue method, which will be described below in detail in conjunction with Figures 2-3 The intelligent home multi-modal dialogue method of the present application is described below.

[0035] Step S11, obtaining device dialogue data.

[0036] In the embodiment, the device conversation data is acquired, including: acquiring device conversation data; wherein the device conversation data includes at least one of audio information, video information, image information and text information; when it is determined that the device conversation data includes audio information, text recognition is performed on the audio information to obtain text information, and the device conversation data is updated using the text information; when it is determined that the device conversation data includes video information, a picture frame sequence is obtained according to the video information, and the device conversation data is updated using the picture frame sequence.

[0037] It should be noted that by acquiring the device conversation data, various original information that may be generated by the device interaction is captured, ensuring that no potential useful data source is missed, and the audio information and video information are converted into structured text and image forms, making subsequent data analysis, understanding, retrieval and application more efficient and possible.

[0038] It should be noted that when the device conversation data includes image information and video information, after obtaining the picture frame sequence according to the video information, the picture frame sequence and the image information are normalized, such as being scaled to a picture of a preset format size. In addition, when the device conversation data includes audio information and text information, after text recognition is performed on the audio information, the text information is normalized, and the normalization includes deleting spaces in the text information, processing split English and special symbols, etc. The specific normalization processing mode can be set according to actual design requirements, which is not limited further herein.

[0039] In step S12, the device conversation data is input into the multi-modal conversation model to obtain a reply prediction result output by the multi-modal conversation model; wherein the multi-modal conversation model is trained according to device historical conversation data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical conversation data; the multi-modal conversation model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to text information in the input device conversation data, and to perform text enhancement on the word features and the part-of-speech features according to the relationship features, to extract context vectors according to the text information, and to decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, to predict a user emotional state according to the device conversation data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0040] In the embodiment, the multi-modal dialogue model comprises: a text enhancement layer configured to extract word features, part-of-speech features corresponding to the word features, and relationship features between adjacent word features from text information in input device dialogue data, perform text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature according to a previous word feature, a part-of-speech feature corresponding to the previous word feature, and a relationship feature between the previous word feature and the subsequent word feature, determine corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtain a text-enhanced word sequence and a text-enhanced part-of-speech sequence; a logical decoding layer configured to extract a context vector from the text information in the device dialogue data, and perform decoding in combination with the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; an emotional expression layer configured to predict a user emotional state according to the text information and image information in the device dialogue data to obtain an emotional prediction state; and a reply generation layer configured to fuse the decoded word sequence and the emotional prediction state, and predict a reply in combination with the fused features to obtain a reply prediction result.

[0041] In other words, the text information of the device dialogue data is input into the multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model, comprising: inputting the device dialogue data into the text enhancement layer to extract word features, part-of-speech features corresponding to the word features, and relationship features between adjacent word features from the text information, perform text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature according to a previous word feature, a part-of-speech feature corresponding to the previous word feature, and a relationship feature between the previous word feature and the subsequent word feature, determine corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtain a text-enhanced word sequence and a text-enhanced part-of-speech sequence; inputting the text information, the text-enhanced word sequence, and the text-enhanced part-of-speech sequence into the logical decoding layer to extract a context vector from the text information in the device dialogue data, and perform decoding in combination with the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; inputting the device dialogue data into the emotional expression layer to predict a user emotional state according to the text information and image information in the device dialogue data to obtain an emotional prediction state; and inputting the decoded word sequence and the emotional prediction state into the reply generation layer to fuse the decoded word sequence and the emotional prediction state, and predict a reply in combination with the fused features to obtain a reply prediction result.

[0042] It should be noted that the text enhancement layer performs deep language feature extraction on the input text information, introducing word features, part-of-speech features, and relational features. This makes the model's understanding of the input text more profound and accurate. Furthermore, it enhances the current word by utilizing the word features and relational features of the previous word, explicitly using explicit language rules such as part-of-speech to strengthen language knowledge constraints and enhance the grammatical meaning and fluency of the text-enhanced word sequence and part-of-speech sequence, making the generated text more fluent. The explicit use of text-enhanced word sequences and part-of-speech sequences ensures that the model follows grammatical rules when generating responses. Combined with context vectors extracted from text information, this ensures the quality of the generated responses. The generated responses are semantically consistent with the dialogue history, producing grammatically and logically sound content. This strengthens explicit knowledge guidance, enhances the logical meaning of the generated decoded word sequences and their focus on the dialogue context, and extracts emotional cues from text and image information to accurately capture the key implicit information of user emotions. This improves implicit knowledge guidance and dynamically balances the semantic and emotional weights of the text, resulting in generated text with rich emotional expression. This enhances the emotional meaning of the generated text and the human touch of the dialogue process. By utilizing implicit emotional knowledge and logical content, the model helps generate more empathetic, appropriate, and resonant responses, making the dialogue experience more natural and human.

[0043] It should be added that the extracted word features are represented as follows: },in, This indicates the first extraction based on text information. Word features This represents the total number of word features extracted based on text information; the extracted part-of-speech features are represented as... ,in, This indicates the first extraction based on text information. The part-of-speech features corresponding to each word feature; the relationship features between extracted adjacent word features are represented as follows: ,in, This indicates the first extraction based on text information. The features of the first word and the first Relationship features between word features.

[0044] In addition, based on the features of the preceding word, the part-of-speech features corresponding to the preceding word, and the relationship between the preceding word and the following word, text enhancement is performed on the features of the following word and the part-of-speech features corresponding to the following word. This includes: concatenating the features of the preceding word and the part-of-speech features corresponding to the preceding word to obtain concatenated features; and generating text-enhanced word features and word features corresponding to the following word based on the word concatenation features and the relationship between the preceding word and the following word.

[0045] It should be noted that for the first word, its word feature and part-of-speech feature are taken as the corresponding text-enhanced word feature and part-of-speech feature, and the subsequent words can be text-enhanced in the manner described above, which will not be repeated here. In addition, the last word feature and its part-of-speech feature are also used to determine the punctuation symbol of the enhanced text representation.

[0046] Further, the text enhancement can be obtained by using a recurrent neural network decoder, and the extracted word features, the part-of-speech features corresponding to the word features, and the relationship features between adjacent word features are input into the decoder to obtain a text-enhanced word sequence and a text-enhanced part-of-speech sequence. For details, please refer to the above, which will not be repeated here.

[0047] In addition, the text-enhanced word features and part-of-speech features are represented as:

[0048]

[0049] wherein, respectively represent the first text-enhanced word features and part-of-speech features; represents the operation of the recurrent neural network decoder; represents the concatenation operation; respectively represent the first word feature and the corresponding part-of-speech feature, i.e., the word feature and the corresponding part-of-speech feature preceding ; represents the total number of word features extracted based on text information.

[0050] Specifically, referring to Figure 2 , the context vector is extracted from the text information in the device dialogue data, and the text-enhanced part-of-speech sequence and the text-enhanced word sequence are decoded to obtain a decoded word sequence, including: extracting the context vector of each text-enhanced word feature corresponding to the text-enhanced word sequence from the text information in the device dialogue data; for each text-enhanced word feature, according to the context vector of each text-enhanced word feature, the text-enhanced part-of-speech sequence and the text-enhanced word sequence, combining the knowledge base, determining the hidden state corresponding to the text-enhanced word feature, and predicting the probability distribution of the next word to obtain a decoded word sequence; wherein the knowledge base is constructed based on the query vector of each word and the knowledge entity of each word.

[0051] It should be noted that by extracting the context vector corresponding to each word feature, the context information is captured to provide a vector representation of each word reflecting its position in the entire dialogue or sentence, which helps the decoder to maintain semantic and logical coherence when generating subsequent words, and the context vector of the word, the text-enhanced part-of-speech feature, and the text-enhanced word feature of the word itself are taken as inputs, and a pre-constructed knowledge base is introduced to further constrain the consistency of the reply content with known facts, thereby determining the corresponding hidden state to guide the generation of the next word, which helps the model to generate replies that conform to grammatical rules and logical order.

[0052] Further, according to the context vector of each text-enhanced word feature, the text-enhanced part-of-speech sequence, and the text-enhanced word sequence, the corresponding hidden state of the text-enhanced word feature is determined in combination with the knowledge base, and the probability distribution of the next word is predicted to obtain the decoding word sequence, including: SA1, when the text-enhanced word feature is the first word in the text-enhanced word sequence, the corresponding context vector is taken as a query vector to search the knowledge base to obtain the corresponding knowledge vector, and the knowledge vector is taken as the corresponding hidden state, and the hidden state is used to predict the probability distribution of the next word; SA2, when the text-enhanced word feature is not the first word in the text-enhanced word sequence, the hidden state of the current word is determined according to the hidden state corresponding to the previous word, the text-enhanced word feature, the text-enhanced part-of-speech feature, the context vector, and the knowledge vector, and the probability distribution of the next word is predicted in combination with the probability distribution of the current word predicted in the previous time, and the hidden state corresponding to the previous word is taken as the query vector of the text-enhanced word feature to search the knowledge base to obtain the knowledge vector corresponding to the text-enhanced word feature; SA3, step SA2 is repeatedly executed until the probability distribution of the last word in the text-enhanced part-of-speech sequence is predicted to obtain the decoding word sequence.

[0053] It should be noted that, for the first word, its context vector is directly used for knowledge base querying to ensure that knowledge retrieval is closely related to the overall dialogue context. The query result is used as the initial hidden state, so that the beginning of the response is strongly influenced by relevant knowledge, which helps to generate starting words that are highly relevant to the current dialogue context and rich in information. For other words, the hidden state of the current word is determined based on the hidden state of the preceding word, the text-enhanced word features, the text-enhanced part-of-speech features, the context vector, and the knowledge vector. By explicitly passing the hidden state of the previous word, the dependency relationship between words is strengthened, which helps to capture long-distance semantic associations, making each generated word more accurate and reasonable. Combined with the probability distribution of the current word predicted in the previous step, the prediction of the current step is more stable and reliable. Step SA2 is repeated until the probability distribution of the last word in the text-enhanced part-of-speech sequence is predicted, ensuring that a complete response sequence matching the length of the input sequence can be generated. This generates an initial response sequence that is both linguistically logical and contains rich knowledge content, laying a solid foundation for subsequent sentiment fusion and final response generation.

[0054] It should be added that the decoded word sequence can be obtained using an explicit recurrent neural network decoder. Accordingly, the hidden state is represented as follows:

[0055]

[0056] in, Indicates the first The hidden state of each word; This represents the operation of the explicit recurrent neural network decoder; Indicates the first The hidden state of each word; Indicates the first Text-enhanced word features of each word; Indicates the first Part-of-speech features; Indicates the first Context vectors; Indicates the first A knowledge vector; Indicates the number of word features; Indicates the first The query vector for the nth word can be obtained through the nth word. The hidden state of each word get; and These represent different mapping matrices, which can be non-shared matrices for parameters. The specific settings can be determined based on actual design requirements or prior experience, and no further limitations are made here. Indicates transpose; Indicates the first The knowledge entity of each word is used to represent the lexical semantic information captured by the corresponding word features and its mapping relationship with the actual word; Indicates the hidden state of the first word; The context vector representing the first word; This represents the query vector for the first word. Indicates based on The knowledge vector of the first word obtained from the knowledge base. It should be added that... The initial computational load can be set according to actual computational needs or prior experience; no further restrictions are imposed here.

[0057] In one optional embodiment, before inputting device dialogue data into the multimodal dialogue model, the process includes: acquiring device historical dialogue data and corresponding answer tags, reference word sequences, and word tags, part-of-speech tags, and sentiment tags for each word in the reference word sequences; wherein, the device historical dialogue data includes historical text information, historical image information, historical video information, and historical audio information corresponding to the target historical time; obtaining historical image frame sequences based on historical video information, and normalizing the historical image frame sequences and historical image information to obtain image training data; performing text recognition on historical audio information to obtain corresponding historical text, and normalizing the historical text and historical text information to obtain text training data; obtaining device dialogue training data based on image training data and text training data; using the device dialogue training data as input data for training, and using word tags, part-of-speech tags, reference word sequences, sentiment tags, and answer tags as tags for training, to train the model to be trained, thereby obtaining the multimodal dialogue model.

[0058] It should be noted that by acquiring historical dialogue data and its labels from the device, decomposing the continuous video stream into discrete image frames, and normalizing it in conjunction with historical image information, the model can be made more robust to different inputs and convergence speed can be accelerated by eliminating differences in brightness, contrast, etc. between different images. Furthermore, audio information is converted into text, ensuring that all inputs are unified into a modality that the model can easily process, facilitating subsequent multimodal fusion. Normalization with historical text ensures the consistency of text data, helping the model learn more stable language representations. This integrates image and text training data from different modalities into a single multimodal training sample pair. Based on real multimodal dialogue data and corresponding comprehensive labels, the model can learn a complete mapping relationship from input to output, learn richer and more robust feature representations, improve overall performance, and transfer knowledge learned from past dialogues to future interactions, enhancing the model's generalization ability and practicality.

[0059] Additionally, refer to Figure 3 After obtaining the text training data, the process includes: dividing the training set according to a preset ratio. and verification set The model is trained using the training set and validated using the validation set. The specific methods are described below and will not be repeated here. , , , , , These represent the image training data in the training set and the validation set, respectively. These represent the text training data in the training set and the validation set, respectively. Indicates the first training set One picture, This indicates the number of images in the training data set. Indicates the first training set A text, This indicates the number of texts in the training data set.

[0060] Specifically, the multimodal dialogue model includes a text enhancement layer, a logic decoding layer, a sentiment expression layer, and a response generation layer. It uses device dialogue training data as input for training, and word tags, part-of-speech tags, reference word sequences, sentiment tags, and response tags as training labels to train the model, resulting in a multimodal dialogue model. This includes: inputting text training data from the device dialogue training data into the text enhancement layer to obtain the text enhancement training word sequence and text enhancement training part-of-speech sequence output by the text enhancement layer; inputting the text training data, text enhancement training word sequence, and text enhancement training part-of-speech sequence into the logic decoding layer to obtain the decoded training word sequence output by the logic decoding layer; and inputting the device dialogue training data into the logic decoding layer to obtain the decoded training word sequence output by the logic decoding layer. Training data is input into the sentiment expression layer to obtain the sentiment training state output by the sentiment expression layer; the decoded training word sequence and the sentiment training state are input into the response generation layer to obtain the response training result output by the response generation layer; a first loss function is constructed based on the text augmentation training word sequence, text augmentation training part-of-speech sequence, word labels, and part-of-speech labels; a second loss function is constructed based on the decoded training word sequence and reference word sequence; a third loss function is constructed based on the sentiment training state and sentiment labels; a fourth loss function is constructed based on the response training result and answer labels; the total loss function is obtained based on the first, second, third, and fourth loss functions, and the training ends upon convergence based on the total loss function.

[0061] It should be noted that the text part in the device dialogue training data is input into the text enhancement layer of the model for language feature extraction and enhancement. By extracting and enhancing word features, part-of-speech features and relationship features, a more rich, semantic and grammatical information representation is generated than the original text. The original text training data, text enhancement training word sequence and text enhancement training part-of-speech sequence are input into the logical decoding layer to combine the context vector and the structured text enhancement features to make a more logical and grammatical prediction, making the decoding process more robust and accurate. The device dialogue training data is input into the sentiment expression layer to combine the text content and the image content to comprehensively judge the emotional state of the user, which helps the model to learn to capture and express emotions and convert complex emotional states into internal representations that the model can process, so that the final generated reply is determined based on language logic (decoded training word sequence) and emotional expression (emotional training state), making the reply more intelligent, natural and human, and through multiple loss functions, the performance of the model on each subtask is more comprehensively evaluated, and through joint optimization, the different parts of the model can coordinate with each other to improve the overall performance.

[0062] In addition, the first loss function is represented as:

[0063]

[0064] wherein, represents the first loss function; represents the total number of word features extracted based on any text in the text training data; respectively represent the first text-enhanced word feature and the part-of-speech feature based on the first text, which can be referred to the above description and will not be repeated here; respectively represent the first word label and the part-of-speech label; represents the cross-entropy loss calculation; represents the activation function softmax loss calculation; represents the word weight matrix; represents the word weight parameter; represents the part-of-speech weight matrix; represents the part-of-speech weight parameter; represents the relationship feature between the first word feature and the second word feature extracted based on all word features of the corresponding text.

[0065] Further, based on the decoded training word sequence and the reference word sequence, a second loss function is constructed, including: aligning the decoded training word sequence and the reference word sequence, and evaluating the word difference loss of each word in the decoded training word sequence and the reference word sequence; determining the true / false label of the corresponding word as false based on the word difference loss being greater than a first preset loss threshold, and determining the true / false label of the corresponding word as true based on the word difference loss being less than or equal to the first preset loss threshold; and constructing the second loss function based on the true / false labels of each word, combined with the hidden state corresponding to each word, the weight matrix used to represent the hidden state of explicit knowledge, and the weight parameters used to represent the hidden state of explicit knowledge.

[0066] It should be noted that the second loss function is expressed as:

[0067]

[0068] in, This represents the second loss function; Indicates the first True or false tags for each word; The weight matrix representing the hidden state of explicit knowledge; Weight parameters representing the hidden state of explicit knowledge; Indicates the first The hidden state of each word can be found in the above text, and will not be repeated here.

[0069] In addition, the emotion training state includes multiple emotion state vectors. Based on the emotion training state and emotion labels, a third loss function is constructed, including: evaluating the emotion difference loss between each emotion state vector and its corresponding emotion label in the emotion training state; determining the emotion true / false label of the corresponding emotion state vector as false based on the emotion difference loss being greater than a second preset loss threshold, and determining the emotion true / false label of the corresponding emotion state vector as true based on the emotion difference loss being less than or equal to the second preset loss threshold; for any emotion state vector, determining the hidden state of the emotion state vector based on the hidden state and memory corresponding to the emotion state vector preceding it, combined with the emotion state vector; and constructing a second loss function based on the emotion true / false labels of each emotion state, combined with the hidden state corresponding to each emotion state, the weight matrix used to represent the hidden state of implicit knowledge, and the weight parameters used to represent the hidden state of implicit knowledge.

[0070] Furthermore, the third loss function is expressed as:

[0071]

[0072] in, Represents the third loss function; Indicates the first The sentiment true / false labels of each sentiment state vector; a weight matrix representing the hidden state of the tacit knowledge; a weight parameter representing the hidden state of the tacit knowledge; representing the operation of the tacit recurrent neural network decoder; representing the hidden state of the emotional state vector; representing the hidden state of the emotional state vector; representing the hidden state of the emotional state vector; representing the memory of the emotional state vector; representing the emotional state vector, which is obtained by mapping the emotional prediction state into the same representation space of the context vector, the emotional prediction state , representing the emotional state vector, then the corresponding emotional state vector , representing the emotional state vector.

[0073] To sum up, the embodiment of the present application inputs the obtained device conversation data into the multi-modal conversation model to extract the relationship features between the word features, the part-of-speech features and the adjacent word features, so as to perform text enhancement on the word features and the part-of-speech features based on the relationship features, enhance the explicit grammatical constraints, make the model more deeply understand the semantics, grammar and context relationship of the input text, make the generated text more fluent, ensure that the generated reply is associated with the entire conversation history by extracting the context vector, decode the part-of-speech and word features after text enhancement, which helps to generate a grammatically correct and logically coherent reply framework, and integrate the knowledge base information to improve the accuracy and knowledge of the reply, and gradually generate words in the decoding process to form a complete reply sequence, then predict the user emotional state according to the text and image information, identify the user's emotional state, so as to generate a reply with more empathy and adaptability, so as to combine the language logic (decoded word sequence) and emotional information (emotional state), and ensure that the generated reply not only meets the logic in content, but also matches the user state in emotional expression.

[0074] The intelligent home multi-modal conversation device provided by the present application is described below, and the intelligent home multi-modal conversation device described below can be correspondingly referred to the intelligent home multi-modal conversation method described above.

[0075] Figure 4 A structural schematic diagram of an intelligent home multi-modal conversation device is shown, the device comprises:

[0076] The data acquisition module 41 acquires device conversation data.

[0077] The reply prediction module 42 inputs the device conversation data into a multi-modal conversation model to obtain a reply prediction result output by the multi-modal conversation model; wherein the multi-modal conversation model is trained according to device historical conversation data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical conversation data; the multi-modal conversation model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to text information in the input device conversation data, and to perform text enhancement on the word features and the part-of-speech features according to the relationship features, to extract context vectors according to the text information, and to decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, to predict a user emotional state according to the device conversation data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0078] In the present embodiment, the data acquisition module 41 comprises: a data acquisition unit that acquires device conversation data; wherein the device conversation data comprises at least one of audio information, video information, image information and text information; a text recognition unit that, when the device conversation data comprises audio information, performs text recognition on the audio information to obtain text information, and updates the device conversation data using the text information; and a picture acquisition unit that, when the device conversation data comprises video information, obtains a picture frame sequence according to the video information, and updates the device conversation data using the picture frame sequence.

[0079] It should be noted that when the device conversation data comprises image information and video information, the data acquisition module 41 further comprises: a normalization processing unit that, after obtaining the picture frame sequence according to the video information, performs normalization processing on the picture frame sequence and the image information, such as scaling to a picture of a preset format size, etc. In addition, when the device conversation data comprises audio information and text information, the normalization processing unit is further used to: after performing text recognition on the audio information, perform normalization processing on the text information, which includes deleting spaces in the text information, processing split English and special symbols, etc. The specific normalization processing mode can be set according to actual design requirements, which is not further limited here.

[0080] In addition, the multi-modal dialogue model comprises: a text enhancement layer configured to extract word features, part-of-speech features corresponding to the word features, and relationship features between adjacent word features from text information in the input device dialogue data, perform text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature according to a previous word feature, a part-of-speech feature corresponding to the previous word feature, and a relationship feature between the previous word feature and the subsequent word feature, determine corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtain a text-enhanced word sequence and a text-enhanced part-of-speech sequence; a logical decoding layer configured to extract a context vector from the text information in the device dialogue data, and perform decoding in combination with the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; an emotional expression layer configured to predict a user emotional state according to the text information and the image information in the device dialogue data to obtain an emotional prediction state; and a reply generation layer configured to fuse the decoded word sequence and the emotional prediction state, and predict a reply in combination with the fused features to obtain a reply prediction result.

[0081] Correspondingly, the reply prediction module 42 comprises: a text enhancement unit configured to input the device dialogue data into the text enhancement layer to extract word features, part-of-speech features corresponding to the word features, and relationship features between adjacent word features from the text information, perform text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature according to a previous word feature, a part-of-speech feature corresponding to the previous word feature, and a relationship feature between the previous word feature and the subsequent word feature, determine corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtain a text-enhanced word sequence and a text-enhanced part-of-speech sequence; a logical decoding unit configured to input the text information, the text-enhanced word sequence, and the text-enhanced part-of-speech sequence into the logical decoding layer to extract a context vector from the text information in the device dialogue data, and perform decoding in combination with the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; an emotional prediction unit configured to input the device dialogue data into the emotional expression layer to predict a user emotional state according to the text information and the image information in the device dialogue data to obtain an emotional prediction state; and a reply generation unit configured to input the decoded word sequence and the emotional prediction state into the reply generation layer to fuse the decoded word sequence and the emotional prediction state, and predict a reply in combination with the fused features to obtain a reply prediction result.

[0082] It should be noted that the text enhancement unit is configured to: concatenate the previous word feature and the part-of-speech feature corresponding to the previous word feature to obtain a concatenated feature; and generate a text-enhanced word feature corresponding to the subsequent word and a text-enhanced part-of-speech feature corresponding to the subsequent word according to the word concatenated feature and the relationship feature between the previous word feature and the subsequent word feature.

[0083] Specifically, the logic decoding unit comprises: a vector extraction subunit configured to extract, according to the text information in the device dialogue data, a context vector corresponding to each text-enhanced word feature of the text-enhanced word sequence; and a decoding subunit configured to, for each text-enhanced word feature, determine a hidden state corresponding to the text-enhanced word feature according to the context vector of the text-enhanced word feature, the text-enhanced part-of-speech sequence and the text-enhanced word sequence, combine a knowledge base, and predict a probability distribution of a next word to obtain a decoding word sequence; wherein the knowledge base is constructed in advance based on the query vectors of the words and the knowledge entities of the words.

[0084] Further, the decoding subunit is configured to: SA1, when the text-enhanced word feature is determined to be the first word in the text-enhanced word sequence, take the corresponding context vector as a query vector to search the knowledge base to obtain a corresponding knowledge vector, take the knowledge vector as the corresponding hidden state, and predict the probability distribution of the next word using the hidden state; SA2, when the text-enhanced word feature is determined to be a non-first word in the text-enhanced word sequence, determine a hidden state of the current word according to the hidden state of the previous word, the text-enhanced word feature, the text-enhanced part-of-speech feature, the context vector and the knowledge vector, predict the probability distribution of the next word in combination with the probability distribution of the current word predicted in the previous time, and take the hidden state corresponding to the previous word as the query vector of the text-enhanced word feature to search the knowledge base to obtain the knowledge vector corresponding to the text-enhanced word feature; and SA3, repeatedly execute step SA2 until the probability distribution of the last word in the text-enhanced part-of-speech sequence is predicted to obtain the decoding word sequence.

[0085] In an optional embodiment, the apparatus further comprises: a training data acquisition module configured to, before inputting the device dialogue data into the multi-modal dialogue model, acquire device historical dialogue data and answer labels corresponding to the device historical dialogue data, reference word sequences and word labels, part-of-speech labels and emotion labels corresponding to the words in the reference word sequences; wherein the device historical dialogue data comprises historical text information, historical image information, historical video information and historical audio information corresponding to a target historical time; a data processing unit configured to, according to the historical video information, obtain a historical picture frame sequence, and perform normalization processing on the historical picture frame sequence and the historical image information to obtain picture training data, and perform text recognition on the historical audio information to obtain corresponding historical text, and perform normalization processing on the historical text and the historical text information to obtain text training data; a data integration unit configured to obtain device dialogue training data according to the picture training data and the text training data; and a training unit configured to train the to-be-trained model by taking the device dialogue training data as input data for training, and taking the word labels, the part-of-speech labels, the reference word sequences, the emotion labels and the answer labels as labels for training to obtain the multi-modal dialogue model.

[0086] Specifically, the training unit comprises: a text enhancement subunit, which inputs text training data in device dialogue training data into a text enhancement layer to obtain a text enhancement training word sequence and a text enhancement training part-of-speech sequence output by the text enhancement layer; a logical decoding subunit, which inputs the text training data, the text enhancement training word sequence and the text enhancement training part-of-speech sequence into a logical decoding layer to obtain a decoding training word sequence output by the logical decoding layer; an emotion prediction subunit, which inputs the device dialogue training data into an emotion expression layer to obtain an emotion training state output by the emotion expression layer; a reply generation subunit, which inputs the decoding training word sequence and the emotion training state into a reply generation layer to obtain a reply training result output by the reply generation layer; a function construction subunit, which constructs a first loss function according to the text enhancement training word sequence, the text enhancement training part-of-speech sequence, a word label and a part-of-speech label, constructs a second loss function according to the decoding training word sequence and a reference word sequence, constructs a third loss function according to the emotion training state and an emotion label, and constructs a fourth loss function according to the reply training result and an answer label; and a training subunit, which obtains a total loss function according to the first loss function, the second loss function, the third loss function and the fourth loss function, and ends the training based on convergence of the total loss function.

[0087] Further, the function construction subunit is configured to: align the decoding training word sequence and the reference word sequence, and evaluate word difference loss of each word in the decoding training word sequence and the reference word sequence; determine a word true-false label of a corresponding word as false based on the word difference loss being greater than a first preset loss threshold, and determine the word true-false label of the corresponding word as true based on the word difference loss being less than or equal to the first preset loss threshold; and construct the second loss function according to the word true-false label of each word, in combination with a hidden state corresponding to each word, a weight matrix used for representing an explicit knowledge hidden state and a weight parameter used for representing the explicit knowledge hidden state.

[0088] In addition, the function construction subunit is further configured to: evaluate emotion difference loss of each emotion state vector in the emotion training state and a corresponding emotion label; determine an emotion true-false label of a corresponding emotion state vector as false based on the emotion difference loss being greater than a second preset loss threshold, and determine the emotion true-false label of the corresponding emotion state vector as true based on the emotion difference loss being less than or equal to the second preset loss threshold; for any emotion state vector, determine a hidden state of the emotion state vector in combination with the emotion state vector according to a hidden state and a memory of an emotion state vector preceding the emotion state vector; and construct the second loss function according to the emotion true-false label of each emotion state, in combination with a hidden state corresponding to each emotion state, a weight matrix used for representing an implicit knowledge hidden state and a weight parameter used for representing the implicit knowledge hidden state.

[0089] In summary, the embodiment of the present application inputs the device conversation data obtained by the data acquisition module into the multi-modal conversation model through the reply prediction module to extract the relationship features between the word features, the part-of-speech features and the adjacent word features, thereby performing text enhancement on the word features and the part-of-speech features based on the relationship features, enhancing the explicit grammatical constraints, enabling the model to more deeply understand the semantics, grammar and context relationship of the input text, making the generated text more fluent, and ensuring that the generated reply is associated with the entire conversation history through the extraction of the context vector, decoding the part-of-speech features and the word features after text enhancement, which helps to generate a reply framework with correct grammar and logical coherence, and integrates knowledge base information to improve the accuracy and knowledge of the reply, and gradually generates words in the decoding process to form a complete reply sequence, and then predicts the user emotional state according to the text and image information, identifies the emotional state of the user, so as to generate a reply with more empathy and adaptability, thereby combining language logic (decoded word sequence) and emotional information (emotional state) to ensure that the generated reply not only meets the logic in content, but also matches the user state in emotional expression.

[0090] Figure 5 An example of an entity structure diagram of an electronic device is shown as Figure 5 The electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can invoke the logic instructions in the memory 530 to execute the smart home multi-modal conversation method, which includes: obtaining device conversation data; inputting the device conversation data into a multi-modal conversation model to obtain a reply prediction result output by the multi-modal conversation model; wherein the multi-modal conversation model is trained according to device historical conversation data and word labels, part-of-speech labels, reference word sequences, emotional labels and answer labels corresponding to the device historical conversation data; the multi-modal conversation model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to the text information in the input device conversation data, and to perform text enhancement on the word features and the part-of-speech features according to the relationship features, to extract context vectors according to the text information, and to decode the part-of-speech features and the word features after text enhancement to obtain a decoded word sequence, to predict the user emotional state according to the device conversation data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0091] In addition, the logical instructions in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0092] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the computer can execute the intelligent home multi-modal dialogue method provided by the above method, the method comprises: obtaining device dialogue data; inputting the device dialogue data into a multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model; wherein the multi-modal dialogue model is trained according to device historical dialogue data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical dialogue data; the multi-modal dialogue model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to the text information in the input device dialogue data, and to perform text enhancement on the word features and the part-of-speech features according to the relationship features, to extract context vectors according to the text information, and to decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, to predict a user emotional state according to the device dialogue data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0093] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the above-provided intelligent home multi-modal dialogue method, which comprises: obtaining device dialogue data; inputting the device dialogue data into a multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model; wherein the multi-modal dialogue model is trained according to device historical dialogue data and word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels corresponding to the device historical dialogue data; the multi-modal dialogue model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to text information in the input device dialogue data, and to perform text enhancement on the word features and the part-of-speech features respectively according to the relationship features, to extract context vectors according to the text information, and to decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, to predict a user emotional state according to the device dialogue data, and to generate a reply prediction result according to the predicted user emotional state and the decoded word sequence.

[0094] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0095] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0096] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A smart home multi-modal dialogue method, characterized in that, The method comprises the following steps: obtaining device dialogue data; inputting the device dialogue data into a multi-modal dialogue model to obtain a reply prediction result output by the multi-modal dialogue model; wherein the multi-modal dialogue model is trained according to device historical dialogue data and corresponding word labels, part-of-speech labels, reference word sequences, sentiment labels and answer labels of the device historical dialogue data; the multi-modal dialogue model is used to extract word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features according to text information in the input device dialogue data, and to perform text enhancement on the word features and the part-of-speech features respectively according to the relationship features, extract context vectors according to the text information, and decode in combination with the text-enhanced part-of-speech features and word features to obtain a decoded word sequence, predict a user emotional state according to the device dialogue data, and generate a reply prediction result according to the predicted user emotional state and the decoded word sequence; the multi-modal dialogue model comprises: a text enhancement layer that extracts word features, part-of-speech features corresponding to the word features and relationship features between adjacent word features from text information in the input device dialogue data, performs text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature according to a previous word feature, a part-of-speech feature corresponding to the previous word feature and a relationship feature between the previous word feature and the subsequent word feature, determines corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtains a text-enhanced word sequence and a text-enhanced part-of-speech sequence; a logical decoding layer that extracts context vectors from text information in the device dialogue data and decodes in combination with the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; an emotional expression layer that predicts a user emotional state according to text information and image information in the device dialogue data to obtain an emotional prediction state; a reply generation layer that fuses the decoded word sequence and the emotional prediction state and predicts a reply in combination with the fused features to obtain a reply prediction result. 2.The smart home multi-modal dialogue method according to claim 1, characterized in that, extracting context vectors from text information in the device dialogue data and decoding in combination with the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence comprises: extracting context vectors corresponding to each text-enhanced word feature of the text-enhanced word sequence from text information in the device dialogue data; for each text-enhanced word feature, determining a hidden state corresponding to the text-enhanced word feature and predicting a probability distribution of a next word in combination with a knowledge base according to the context vectors of each text-enhanced word feature, the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; wherein the knowledge base is constructed in advance based on query vectors of each word and knowledge entities of each word. 3.The smart home multi-modal dialogue method according to claim 2, characterized in that, determining a hidden state corresponding to each text-enhanced word feature and predicting a probability distribution of a next word in combination with a knowledge base according to the context vectors of each text-enhanced word feature, the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence comprises: SA1, when determining that the text-enhanced word feature is the first word in the text-enhanced word sequence, taking the corresponding context vector as a query vector to search the knowledge base to obtain a corresponding knowledge vector, and taking the knowledge vector as a corresponding hidden state, and using the hidden state to predict a probability distribution of a next word; SA2, when determining that the text-enhanced word feature is not the first word in the text-enhanced word sequence, determining a hidden state of a current word according to a hidden state corresponding to a previous word, the text-enhanced word feature, a text-enhanced part-of-speech feature, a context vector and a knowledge vector, and combining a probability distribution of the current word predicted in a previous time to predict a probability distribution of a next word, and taking the hidden state corresponding to the previous word as a query vector of the text-enhanced word feature to search the knowledge base to obtain a knowledge vector corresponding to the text-enhanced word feature; SA3, repeatedly performing step SA2 until a probability distribution of a last word in the text-enhanced part-of-speech sequence is predicted to obtain a decoded word sequence. 4.The smart home multi-modal dialogue method according to claim 1, characterized in that, Before inputting the device dialogue data into the multi-modal dialogue model, comprising: obtaining device historical dialogue data and answer labels corresponding to the device historical dialogue data, reference word sequences and word labels, part-of-speech labels and emotion labels corresponding to words in the reference word sequences; wherein the device historical dialogue data comprises historical text information, historical image information, historical video information and historical audio information corresponding to a target historical time; obtaining a historical picture frame sequence according to the historical video information, and performing normalization processing on the historical picture frame sequence and the historical image information to obtain picture training data; performing text recognition on the historical audio information to obtain corresponding historical text, and performing normalization processing on the historical text and the historical text information to obtain text training data; obtaining device dialogue training data according to the picture training data and the text training data; training a to-be-trained model by taking the device dialogue training data as input data for training and taking the word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels as labels for training to obtain a multi-modal dialogue model. 5.The smart home multi-modal dialogue method according to claim 4, characterized in that, The multi-modal dialogue model comprises a text enhancement layer, a logical decoding layer, an emotion expression layer and a reply generation layer; a to-be-trained model is trained by taking the device dialogue training data as input data for training and taking the word labels, part-of-speech labels, reference word sequences, emotion labels and answer labels as labels for training to obtain a multi-modal dialogue model, comprising: inputting text training data in the device dialogue training data into the text enhancement layer to obtain text-enhanced training word sequences and text-enhanced training part-of-speech sequences output by the text enhancement layer; inputting the text training data, the text-enhanced training word sequences and the text-enhanced training part-of-speech sequences into the logical decoding layer to obtain a decoded training word sequence output by the logical decoding layer; inputting the device dialogue training data into the emotion expression layer to obtain an emotion training state output by the emotion expression layer; inputting the decoded training word sequence and the sentiment training state into the reply generation layer to obtain a reply training result output by the reply generation layer; constructing a first loss function according to the text-enhanced training word sequence, the text-enhanced training part-of-speech sequence, the word label, and the part-of-speech label; constructing a second loss function according to the decoded training word sequence and the reference word sequence; constructing a third loss function according to the sentiment training state and the sentiment label; constructing a fourth loss function according to the reply training result and the answer label; obtaining a total loss function according to the first loss function, the second loss function, the third loss function, and the fourth loss function, and ending the training based on convergence of the total loss function. 6.The smart home multi-modal dialogue method according to claim 1, characterized in that, obtaining device conversation data, including: obtaining device conversation data; wherein the device conversation data includes at least one of audio information, video information, image information, and text information; when the device conversation data includes audio information, performing text recognition on the audio information to obtain text information, and updating the device conversation data using the text information; when the device conversation data includes video information, obtaining a picture frame sequence according to the video information, and updating the device conversation data using the picture frame sequence.

7. A smart home multi-modal dialogue device, characterized in that, including: a data acquisition module that obtains device conversation data; a reply prediction module that inputs the device conversation data into a multi-modal conversation model to obtain a reply prediction result output by the multi-modal conversation model; wherein the multi-modal conversation model is trained according to device historical conversation data and word labels, part-of-speech labels, reference word sequences, sentiment labels, and answer labels corresponding to the device historical conversation data; the multi-modal conversation model is configured to extract word features, part-of-speech features corresponding to the word features, and relationship features between adjacent word features from text information in input device conversation data, perform text enhancement on the word features and the part-of-speech features based on the relationship features, extract context vectors from the text information, and decode the part-of-speech features and the word features after text enhancement to obtain a decoded word sequence, predict a user sentiment state based on the device conversation data, and generate a reply prediction result based on the predicted user sentiment state and the decoded word sequence; the multi-modal conversation model includes: a text enhancement layer that extracts word features, part-of-speech features corresponding to the word features, and relationship features between adjacent word features from text information in input device conversation data, performs text enhancement on a subsequent word feature and a part-of-speech feature corresponding to the subsequent word feature based on a previous word feature, a part-of-speech feature corresponding to the previous word feature, and relationship features between the previous word feature and the subsequent word feature, determines corresponding text-enhanced word features and text-enhanced part-of-speech features, and obtains a text-enhanced word sequence and a text-enhanced part-of-speech sequence; a logical decoding layer that extracts context vectors from text information in the device conversation data, and decodes the text-enhanced part-of-speech sequence and the text-enhanced word sequence to obtain a decoded word sequence; An emotion expression layer is configured to predict a user emotion state according to text information and image information in the device dialogue data, and obtain an emotion prediction state; A reply generation layer is configured to fuse the decoded word sequence and the emotion prediction state, and predict a reply in combination with a fusion feature, and obtain a reply prediction result.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the smart home multi-modal dialogue method according to any one of claims 1 to 6 when executing the computer program. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the smart home multi-modal dialogue method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • BERT-based smart home use event extraction method

    CN117390175A

  • Multi-modal emotion understanding method based on cross-modal semantic alignment and interactive learning

    CN119475214A