Multimodal Speech Interaction Method, Electronic Device, and Storage Medium Based on Large Language Model

Through the multimodal voice interaction method based on large models, the problems of multilingual communication and multimodal information fusion in the existing technology are solved, and efficient and intelligent voice interaction is achieved.

CN119559946BActive Publication Date: 2025-06-13SHENZHEN BEIBO INTELLIGENT TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510026491.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-13
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing voice interaction technologies are unable to effectively handle multilingual communication and multimodal information fusion, resulting in limited interaction fluency and usability.

Method used

A multimodal speech interaction method based on large models is adopted to generate personalized dialogue strategies through speech recognition, language translation, semantic analysis and multimodal knowledge graph construction, and perform speech synthesis to realize multilingual interaction and multimodal information fusion.

Benefits of technology

It realizes the convenience of cross-language interaction and the effective utilization of multimodal information, and improves the fluency, intelligence and accuracy of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559946B_ABST
    Figure CN119559946B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal speech interaction method, an electronic device and a storage medium based on a large model, including: performing speech recognition on speech data in a first language input by a user to obtain a first language text; performing language translation on the first language text to translate it into a second language text; performing semantic analysis on the second language text to obtain a semantic analysis result; constructing a multi-modal knowledge graph based on a preset large model for the semantic analysis result to obtain an enhanced semantic understanding result; wherein the multi-modal knowledge graph integrates multi-modal information related to the second language text; generating a dialogue strategy for the enhanced semantic understanding result; and performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language. In the present invention, multi-language interaction is realized, and at the same time, various modal information is integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly to a multimodal speech interaction method, an electronic device, and a storage medium based on a large model. Background Art

[0002] Speech interaction technology has become one of the key technologies in the field of human-computer interaction, aiming to provide users with a more convenient, natural, and efficient interaction experience.

[0003] In traditional speech interaction systems, it mainly relies on speech recognition of a single language and simple text processing technologies to understand the user's intention and generate a response. However, this method has many limitations. On the one hand, with the increasing frequency of cross-border communication and multicultural integration, users often need to communicate in different languages in actual usage scenarios. A single-language speech interaction system cannot meet this demand, resulting in language barriers for users when interacting with devices, which greatly affects the fluency and usability of the interaction.

[0004] On the other hand, existing speech interaction technologies usually only process based on speech text information and lack effective integration and utilization of other relevant modal information. For example, in many actual situations, multimodal information such as images and videos related to the speech content can provide rich context and supplementary information for accurately understanding the user's intention, but traditional systems are difficult to incorporate this multimodal information into the interaction process, resulting in an insufficiently intelligent and accurate interaction process that cannot meet the requirements of complex and changing real-world scenarios. Summary of the Invention

[0005] The main object of the present invention is to provide a multimodal speech interaction method, an electronic device, and a storage medium based on a large model, aiming to overcome the defects of the current inability to interact in multiple languages and the inability to integrate multiple modal information.

[0006] To achieve the above object, the present invention provides a multimodal speech interaction method based on a large model, including the following steps:

[0007] Perform speech recognition on the speech data of the first language input by the user to obtain the first language text;

[0008] Perform language translation on the first language text to translate it into the second language text; perform semantic analysis on the second language text to obtain a semantic analysis result;

[0009] Based on a preset large model, construct a multimodal knowledge graph for the semantic analysis result to obtain an enhanced semantic understanding result; wherein, the multimodal knowledge graph integrates multimodal information related to the second language text;

[0010] A dialogue strategy for generating the enhanced semantic understanding result;

[0011] Based on the dialogue strategy, perform speech synthesis to generate response speech data expressed in the first language.

[0012] Further, the dialogue strategy for generating the enhanced semantic understanding result includes:

[0013] Based on a preset intention classification model, divide the enhanced semantic understanding result into multiple intention categories; wherein, the intention classification model is obtained through pre-training;

[0014] For each intention category, match the corresponding dialogue strategy template from a pre-stored dialogue strategy template library; the dialogue strategy template library is pre-constructed based on multi-language and multi-scenario dialogue data;

[0015] Using natural language generation technology, fill in and adjust the content of the matched dialogue strategy template according to the specific content in the enhanced semantic understanding result to generate a personalized dialogue strategy.

[0016] Further, based on a preset large model, construct a multi-modal knowledge graph for the semantic analysis result to obtain an enhanced semantic understanding result, including:

[0017] Encode the semantic analysis result based on the text encoder in the large model to obtain a text vector;

[0018] For multi-modal information related to the second language text, use the corresponding image encoder and video encoder to extract and encode the image data and video data respectively to obtain an image feature vector and a video feature vector;

[0019] Fuse the image feature vector and the video feature vector into the text vector to construct an initial knowledge graph; wherein, the text vector is the core node, and the fused multi-modal features are used as peripheral nodes and edges to form an initial knowledge graph;

[0020] Optimize the initial knowledge graph based on a semantic relationship model to obtain an enhanced semantic understanding result; wherein, the semantic relationship model is unsupervised learning based on a multi-modal corpus, can identify and supplement semantic relationships in the knowledge graph, and improve the structure and content of the knowledge graph by continuously iteratively adjusting the connection relationships and semantic annotations between nodes.

[0021] Further, fusing the image feature vector and the video feature vector into the text vector includes:

[0022] Through a cross-modal attention mechanism, the text vector is interacted with the image feature vector and the video feature vector to calculate the correlation weights between different modalities; wherein, the cross-modal attention mechanism is based on a multi-head attention architecture, and captures the correlation information between different modalities by parallelly calculating multiple attention heads;

[0023] According to the calculated correlation weights, the image feature vector and the video feature vector are fused into the text vector.

[0024] Further, after performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language, it includes:

[0025] Generating a speech data encryption mode based on the codes of the first language and the second language;

[0026] Generating a communication key based on the semantic analysis result;

[0027] After encrypting the reply speech data with the communication key based on the speech data encryption mode, it is fed back to the user's terminal.

[0028] Further, the generating a speech data encryption mode based on the codes of the first language and the second language includes:

[0029] Based on a language code mapping table, mapping the codes of the first language and the second language to unique first digital identifiers and second digital identifiers respectively;

[0030] Respectively calculating a first difference and a second difference between the first digital identifier and the second digital identifier and a preset encryption base value;

[0031] Encrypting the sum value of the first difference and the second difference based on a preset encryption algorithm to obtain an encrypted feature value;

[0032] Performing binary decomposition on the encrypted feature value, and analyzing it in combination with the frequency and duration of the reply speech data, and determining the encryption segmentation method and local encryption area through a dynamic programming algorithm.

[0033] Further, the generating a communication key based on the semantic analysis result includes:

[0034] Extracting the lexical feature vector, syntactic feature vector and semantic dependency relationship feature vector in the semantic analysis result;

[0035] Constructing a graph structure based on each feature vector; wherein, the lexical feature vector is used as a node, and the syntactic feature vector and the semantic dependency relationship feature vector are used to determine the edge weights;

[0036] Perform a depth - first traversal on the graph structure and map the obtained node sequence to a numerical sequence;

[0037] Generate a polynomial curve using the numerical sequence, generate a key seed based on the coefficients and control point information of the polynomial curve; Hash the key seed using a hash function to obtain a hash value;

[0038] Use the hash value as input and generate a communication key through a key - stream generation algorithm based on a linear feedback shift register.

[0039] Furthermore, generating a communication key based on the semantic analysis result includes:

[0040] Encode the semantic analysis result to obtain encoded characters;

[0041] Insert the encoded characters into each node of the graph structure one by one in sequence, and assign values to the edges between adjacent nodes according to the relationship between the characters on adjacent nodes to obtain a character graph structure;

[0042] Find all edges with different assignments as target edges; Connect the mid - points of each target edge in sequence to obtain multiple target connections;

[0043] Find the nodes in the character graph structure that satisfy a preset relationship with each target connection as target nodes;

[0044] Combine the characters on the target nodes in sequence as the key seed, hash the key seed using a hash function to obtain a hash value;

[0045] Use the hash value as input and generate a communication key through a key - stream generation algorithm based on a linear feedback shift register.

[0046] The present invention also provides a multi - modal voice interaction device based on a large model, including:

[0047] An identification module for performing speech recognition on the speech data of the first language input by the user to obtain the first - language text;

[0048] An analysis module for translating the first - language text into a second - language text; performing semantic analysis on the second - language text to obtain a semantic analysis result;

[0049] A construction module for constructing a multi - modal knowledge graph based on a preset large model for the semantic analysis result to obtain an enhanced semantic understanding result; wherein, the multi - modal knowledge graph integrates multi - modal information related to the second - language text;

[0050] A generation module for generating a dialogue strategy for the enhanced semantic understanding result;

[0051] A synthesis module for performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language.

[0052] The present invention also provides an electronic device including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0053] The present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0054] The multi-modal voice interaction method, electronic device, and storage medium based on a large model provided by the present invention include: performing speech recognition on the speech data in the first language input by a user to obtain a first-language text; performing language translation on the first-language text to translate it into a second-language text; performing semantic analysis on the second-language text to obtain a semantic analysis result; constructing a multi-modal knowledge graph based on a preset large model for the semantic analysis result to obtain an enhanced semantic understanding result, where the multi-modal knowledge graph integrates multi-modal information related to the second-language text; generating a dialogue strategy for the enhanced semantic understanding result; performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language. In the present invention, speech recognition and language translation are performed on the speech data in the first language input by the user to obtain a second-language text, and then semantic analysis is performed to obtain a semantic analysis result, and a multi-modal knowledge graph is constructed, and finally reply speech data is generated. Multi-language interaction is achieved while integrating various modal information. Description of the Drawings

[0055] Figure 1 is a schematic diagram of the steps of the multi-modal voice interaction method based on a large model in an embodiment of the present invention;

[0056] Figure 2 is a block diagram of the structure of the multi-modal voice interaction device in an embodiment of the present invention;

[0057] Figure 3 is a schematic block diagram of the structure of an electronic device in an embodiment of the present invention.

[0058] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0059] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0060] Referring Figure 1 , in one embodiment of the present invention, a multi-modal voice interaction method based on a large model is provided, including the following steps:

[0061] Step S1, perform speech recognition on the speech data of the first language input by the user to obtain the first language text;

[0062] Step S2, perform language translation on the first language text to translate it into a second language text; perform semantic analysis on the second language text to obtain a semantic analysis result;

[0063] Step S3, based on a preset large model, construct a multi-modal knowledge graph for the semantic analysis result to obtain an enhanced semantic understanding result; wherein, the multi-modal knowledge graph integrates multi-modal information related to the second language text;

[0064] Step S4, generate a dialogue strategy for the enhanced semantic understanding result;

[0065] Step S5, perform speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language.

[0066] In this embodiment, as described in step S1 above, the speech information spoken by the user in the first language is converted into a corresponding text form for subsequent in-depth processing and analysis. It utilizes speech recognition technology, which is usually based on components such as an acoustic model and a language model. The acoustic model models and analyzes the acoustic features (such as phoneme features, etc.) in the speech signal to identify the phoneme sequence contained in the speech; the language model, based on the grammar, vocabulary and other rules of the language, combines the identified phoneme sequence to infer the most likely word combination, and finally forms the complete first language text.

[0067] As described in step S2 above, considering that in a cross-language communication scenario, it is necessary to convert the first language text input by the user into the target second language text in order to meet the interaction needs of users with different language backgrounds. This is often done with the help of machine translation technology, such as methods based on statistical machine translation or neural machine translation. Neural machine translation uses deep neural networks, such as the Transformer architecture, to learn a large amount of bilingual parallel corpus, which can automatically capture the mapping relationship between the two languages, and convert the vocabulary, grammatical structure and other information in the first language text into the corresponding expression form in the second language according to the learned pattern, so that the system can communicate and interact effectively with more users who use different languages.

[0068] Semantic analysis is to deeply understand the actual meaning conveyed by the translated second language text and to mine the semantic relationship, theme, intention and other information. This process may use technical means such as lexical analysis, syntactic analysis and semantic role labeling in natural language processing. Lexical analysis will perform operations such as part-of-speech tagging on words in the text; syntactic analysis constructs the grammatical structure tree of the sentence to clarify the relationship between the components; semantic role labeling further determines the semantic role played by each entity in the text, and combines this information to obtain semantic analysis results, such as determining whether a sentence is asking for information, expressing a request or stating a fact. Accurate semantic analysis is the key basis for the subsequent construction of multimodal knowledge graphs and the generation of appropriate dialogue strategies. Only by understanding the semantics of the text can we give a response that fits the user's intention in the interaction.

[0069] As described in step S3 above, the above-mentioned large model, with its powerful parameter scale and learning ability, can deeply mine the semantic analysis results and integrate related multimodal information (such as pictures, videos, audio and other modal contents) to construct a knowledge graph. During the construction process, the large model will first extract features of information of different modalities, such as extracting image features using a convolutional neural network, extracting audio features using an audio processing algorithm, etc., and then find the connection between text semantics and other modal information through a cross-modal association learning mechanism, and structure these associations in the form of a knowledge graph. For example, if the second language text mentions "giant panda", the large model can be associated with multimodal information such as stored pictures of giant pandas and videos introducing the living habits of giant pandas, and construct knowledge graph nodes around the theme of "giant pandas" and the relationship edges between them, thereby enhancing the understanding of text semantics. In this embodiment, the limitation of traditional voice interaction relying only on text information is broken, and the dimension of semantic understanding is enriched by making full use of multimodal information to make the system more accurate and comprehensive in grasping the user's intentions, thereby improving the quality and intelligence of the entire voice interaction.

[0070] As described in step S4 above, based on the semantic understanding result enhanced by the multi-modal knowledge graph, an appropriate dialogue strategy is formulated. For example, if the semantic understanding result shows that the user is asking about the function of a certain product, dialogue strategies such as introducing the product function in detail and providing relevant cases will be generated according to the pre-stored rules or the effective response methods learned from past similar scenarios to guide the smooth progress of the dialogue and meet the user's information needs.

[0071] As described in step S5 above, the text content corresponding to the previously generated dialogue strategy is converted back into a voice form through text-to-speech technology and expressed in the first language initially input by the user. The text-to-speech technology will model the pronunciation of words, phrases, etc. in the text according to the phonetic rules of the language, and generate natural and fluent voice signals through methods such as splicing voice units or parametric synthesis. The complete closed-loop from the user's voice input to the system's voice response is completed, and the interaction result is presented to the user in the form of voice, which conforms to the natural and convenient characteristics of voice interaction, allowing the user to intuitively obtain the system's response content and realizing efficient human-computer interaction.

[0072] In the present invention, the speech data of the first language input by the user is subjected to speech recognition and language translation to obtain the second language text, and then semantic analysis is performed to obtain the semantic analysis result, and a multi-modal knowledge graph is constructed, and finally the response speech data is generated. It realizes interaction in multiple languages and integrates multiple modal information at the same time.

[0073] In one embodiment, the dialogue strategy for generating the enhanced semantic understanding result includes:

[0074] Based on a preset intention classification model, the enhanced semantic understanding result is divided into multiple intention categories; wherein, the intention classification model is obtained through pre-training;

[0075] For each intention category, a corresponding dialogue strategy template is matched from the pre-stored dialogue strategy template library; the dialogue strategy template library is pre-constructed based on multi-language and multi-scenario dialogue data;

[0076] Using natural language generation technology, according to the specific content in the enhanced semantic understanding result, the matched dialogue strategy template is filled and adjusted to generate a personalized dialogue strategy.

[0077] In this embodiment, the above-mentioned intent classification is to extract the core intent of the user from the semantic understanding results after the multimodal knowledge graph is enhanced, so as to accurately match the corresponding dialogue strategy in the future. The intent classification model is usually built based on machine learning or deep learning algorithms, such as support vector machines, convolutional neural networks (CNN) or recurrent neural networks (RNNS) in deep learning and their variants (such as LSTM, GRU, etc.). During the training phase, a large amount of multilingual and multi-scenario text data with intent labels will be collected as training samples. These samples cover various possible user expression intentions, such as asking for information, seeking help, expressing opinions, issuing instructions and other different categories. The model learns how to distinguish different intent categories by learning the mapping relationship between the vocabulary, grammar, semantics and other features of the text in these samples and the corresponding intent labels.

[0078] When faced with new enhanced semantic understanding results, the model can extract key features based on the learned patterns and classify them into corresponding intent categories.

[0079] The construction of the above-mentioned dialogue strategy template library is to reserve a general dialogue response framework for different intentions, different languages ​​and different scenarios in advance. During the construction process, a large amount of real dialogue data is collected, covering expressions in multiple languages ​​and various practical application scenarios (such as shopping, travel, social networking, etc.), and these data are analyzed and sorted to extract common and representative dialogue structures to form dialogue strategy templates. Each template corresponds to a specific intent category. After the intent category to which the enhanced semantic understanding result belongs is determined by the intent classification model, the template that matches it can be searched and matched from this pre-built library. In this embodiment, the efficiency and standardization of dialogue strategy generation are greatly improved. By using the pre-built template library, the tedious process of generating dialogue strategies from scratch each time is avoided, while ensuring that the response conforms to the common effective dialogue mode in structure and logic, laying the foundation for generating high-quality dialogue strategies.

[0080] Natural language generation technology aims to enable computers to generate text content that is fluent and accurate in human language habits. In this step, it will improve the matched dialogue strategy template based on the specific details in the enhanced semantic understanding results. For example, there may be some placeholders in the template that need to be filled according to the specific semantic content.

[0081] If a specific product, such as "a certain brand of smartphone", is mentioned in the enhanced semantic understanding result, and there are placeholders for product brand and type in the matching template, the natural language generation technology will accurately fill in the corresponding information. At the same time, according to the context and semantic relationships, it will appropriately adjust the wording, sentence order, etc. of the entire template to make it more natural and appropriate, and finally generate a personalized dialogue strategy, such as "This certain brand of smartphone has functions such as a high-definition screen and a large-capacity battery, which can help you use it for a long time and obtain a clear visual experience." This makes the dialogue strategy fit the user's specific expression and actual needs, achieving personalized responses. It avoids the uniformity of the response content, allowing users to feel that the system truly understands their intentions and gives targeted responses, thereby enhancing the intelligence and user satisfaction of the voice interaction and making the entire interaction process more natural and smooth.

[0082] In one embodiment, based on a preset large model, a multi-modal knowledge graph is constructed for the semantic analysis result to obtain an enhanced semantic understanding result, including:

[0083] Encoding the semantic analysis result based on the text encoder in the large model to obtain a text vector;

[0084] For the multi-modal information related to the second language text, use the corresponding image encoder and video encoder to respectively extract and encode the image data and video data to obtain an image feature vector and a video feature vector;

[0085] Fuse the image feature vector and the video feature vector into the text vector to construct an initial knowledge graph; wherein, the text vector is the core node, and the fused multi-modal features are used as peripheral nodes and edges to form an initial knowledge graph;

[0086] Optimizing the initial knowledge graph based on a semantic relationship model to obtain an enhanced semantic understanding result; wherein, the semantic relationship model is unsupervised learning based on a multi-modal corpus, which can identify and supplement the semantic relationships in the knowledge graph, and continuously iteratively adjust the connection relationships and semantic annotations between nodes to improve the structure and content of the knowledge graph.

[0087] In this embodiment, the above text encoder plays a key role in vectorizing the information in the form of text of the semantic analysis result in the large model. Its principle is usually based on the word embedding technology in deep learning and more complex sequence encoding mechanisms, such as the encoding layer in the Transformer architecture, etc.

[0088] Word embedding technology maps each word in the text into a low-dimensional vector space, making words with similar semantics closer in this vector space, so as to capture the semantic features of words. For the entire text sequence, the encoder further considers the order relationship between words as well as the combination of grammar and semantics. Through a multi-layer neural network structure, the text is deeply encoded, and finally the semantic analysis result is transformed into a text vector that can represent its overall semantics. The acquisition of the text vector is the basis for subsequent multi-modal fusion and the construction of knowledge graphs. After converting the text information into vector form, it is convenient for the computer to perform numerical processing and analysis, enabling it to interact with vector representations of other modalities in the same mathematical space, thus laying the foundation for integrating multi-modal information. It is the first step in realizing multi-modal fusion and directly affects the accuracy and effectiveness of the construction of the entire multi-modal knowledge graph.

[0089] Image encoders and video encoders are designed to extract key features from image and video data and transform them into vector representation forms so that they can be fused with text vectors. For image encoders, they are often based on the convolutional neural network (CNN) architecture. Through components such as convolutional layers and pooling layers, CNN can automatically learn visual features such as textures, shapes, and colors in images. As the number of network layers deepens, more abstract and representative high-level features are gradually extracted, and finally the entire image is compressed into an image feature vector.

[0090] Video encoders need to consider the characteristics of video data in the time dimension. In addition to using a structure similar to CNN to extract features from each frame of the image, they also combine recurrent neural networks (RNN) or their variants (such as LSTM, GRU, etc.) to capture the temporal relationship between frames, thereby extracting the overall features of the video and obtaining a video feature vector.

[0091] By extracting and encoding image and video data, the conversion of different modal information into a unified vector representation form is achieved, enabling multi-modal information to be effectively fused with text information in subsequent steps. Only by converting the information of each modality into a vector form that can be compared and calculated with each other can the modal barriers be broken and the multi-modal information be integrated together, providing the necessary conditions for constructing a comprehensive and rich multi-modal knowledge graph and helping to understand the content related to the second language text more comprehensively.

[0092] Fuse the feature vectors of different modalities and organize them in a structured manner in the form of a knowledge graph. Some specific fusion strategies are usually adopted in the fusion process. For example, in the method based on the Attention Mechanism, the attention mechanism can dynamically calculate the correlation weights between the feature vectors of different modalities. According to these weights, the key information in the image feature vector and the video feature vector is fused onto the text vector, highlighting the important associated parts between modalities.

[0093] When constructing the initial knowledge graph, the text vector is used as the core node because the text carries the core content of semantics, while the fused multi-modal feature vectors are used as peripheral nodes and edges to represent the multi-modal information such as images and videos related to the text and their relationships with each other. The first integration of multi-modal information is achieved. By constructing the initial knowledge graph, the originally scattered multi-modal information such as text, images, and videos is associated in a structured way, initially forming a knowledge system that can reflect the multi-modal semantic relationships, providing a basic framework for further optimizing and improving the knowledge graph in the future, so as to obtain enhanced semantic understanding results, and helping to explore the supplementary and strengthening effects of multi-modal information on semantic understanding.

[0094] The purpose of the above semantic relationship model is to deeply optimize the initially constructed knowledge graph so that it can more accurately and comprehensively reflect the semantic relationships between multi-modal information. This model performs unsupervised learning based on a multi-modal corpus, which means it automatically learns the semantic association patterns in a large amount of multi-modal data including text, images, videos, etc., without manually annotating specific relationship types and other information. By using the semantic relationship model to improve the knowledge graph, more hidden and deep multi-modal semantic relationships can be mined, making up for the possible deficiencies in the initial construction process. As a result, the finally obtained knowledge graph can provide richer and more accurate information for semantic understanding, achieving the goal of enhanced semantic understanding, enabling the system to more accurately grasp the user's intention and understand the text meaning based on the multi-modal knowledge graph, and thus improving the quality and effect of the entire multi-modal speech interaction.

[0095] In one embodiment, fusing the image feature vector and the video feature vector into the text vector includes:

[0096] Through the cross-modal attention mechanism, the text vector interacts with the image feature vector and the video feature vector to calculate the correlation weights between different modalities; wherein, the cross-modal attention mechanism is based on the multi-head attention architecture, and captures the correlation information between different modalities by calculating multiple attention heads in parallel;

[0097] According to the calculated correlation weights, fuse the image feature vector and the video feature vector into the text vector.

[0098] In this embodiment, in multimodal information processing, text, images, and videos carry information in different dimensions and exist in vector form respectively. To effectively integrate this information from different modalities, a mechanism is needed to measure the degree of association between them. The purpose of the cross-modal attention mechanism is to dynamically determine the importance weights of the association between the text vector and the image feature vector and the video feature vector, so that the multimodal information can be reasonably fused based on these weights subsequently.

[0099] The above-mentioned multi-head attention architecture is an advanced and efficient design. It captures the association information between different modalities by computing multiple attention heads in parallel. Each attention head can be regarded as an independent "attention perspective" to analyze the correlation between the text vector and other modality vectors from different angles. Taking the multi-head attention mechanism in the common Transformer architecture as an example, it first performs linear transformations on the input text vector, image feature vector, and video feature vector respectively to obtain the corresponding query, key, and value representations. For each attention head, the attention score is obtained by calculating the similarity between the query and the key (usually using methods such as dot product), and this score represents the degree of association between different modalities under this attention head. Then, through operations such as normalizing the attention score and multiplying and accumulating it with the corresponding value, the feature representation after being attended to and weighted by this attention head is obtained. Multiple attention heads execute such operations in parallel, each capturing the association information between modalities at different levels and from different angles. Finally, the results of these different attention heads are concatenated or fused, etc., to comprehensively obtain the comprehensive and rich modality association weight information. For example, for a text vector describing "beach scenery", as well as the corresponding image feature vector showing the beach scene and the video feature vector recording beach activities, different attention heads respectively focus on the association between the "waves" mentioned in the text and the wave form in the image and the wave undulation dynamics in the video, or the association between the "beach chair" in the text and the placement of the beach chair in the image and the scene of people using the beach chair in the video, etc., and then determine the corresponding association weights.

[0100] In this embodiment, it is possible to accurately calculate the degree of association between different modalities in terms of semantics, content, etc., avoiding the problem of simply fusing multimodal information while ignoring their internal connections. By multiple attention heads mining the association information from multiple angles, it more comprehensively reflects the complex mutual relationship between modalities, providing a reliable data basis for subsequent high-quality multimodal fusion. It is a key prerequisite step to effectively fuse the image and video multimodal feature vectors into the text vector, directly determining whether the fusion effect can fully utilize the advantages of multimodal information and improve the accuracy and richness of semantic understanding.

[0101] After obtaining the correlation weights between the text vector and the image feature vector and the video feature vector, it is necessary to fuse the feature vectors of the image and the video into the text vector according to these weights to truly achieve the integration of multimodal information. The specific fusion operation is usually to perform weighted summation on each modal vector according to the weights. By performing such operations on each dimension and fusing according to the calculated correlation weights, the key information in the image and video feature vectors is incorporated into the text vector, making the text vector no longer only contain the semantic information of the text itself, but also incorporate the related image and video multimodal information, forming a new vector representation that synthesizes multimodal features. This step is the key final link in the entire multimodal fusion process. By fusing the features of different modalities according to reasonable correlation weights, the barriers between modalities are successfully broken, and multimodal information is gathered together to construct a unified vector representation that can comprehensively reflect the relevant content of text, image, and video. This fused vector serves as the basis for subsequent further processing such as constructing a knowledge graph, and plays a crucial role in enriching semantic understanding and enhancing the content grasping ability of the multimodal-based speech interaction system, enabling the system to understand and process information related to user input from multiple perspectives and dimensions, thereby improving the quality and intelligence level of the entire speech interaction.

[0102] In one embodiment, after performing speech synthesis based on the dialogue strategy and generating reply speech data expressed in the first language, it includes:

[0103] Generating a speech data encryption mode based on the codes of the first language and the second language;

[0104] Generating a communication key based on the semantic analysis result;

[0105] After encrypting the reply speech data with the communication key based on the speech data encryption mode, feeding it back to the user's terminal.

[0106] In this embodiment, during the multimodal speech interaction process, the generated reply speech data may contain user-sensitive information or involve privacy content. To prevent this data from being illegally obtained, tampered with, etc. during transmission, it needs to be encrypted. Generating an encryption mode based on the codes of the first language and the second language takes into account the characteristics of different languages (such as grammar structures, vocabulary features, etc.) and the language conversion situations involved in speech interaction, and customizes a dedicated encryption method for speech data to improve the pertinence and security of encryption.

[0107] First, the codes of the first language and the second language are usually specific encodings or identifiers used to uniquely identify these two languages. By using these codes and combining pre-set rules or algorithms, the encryption mode is determined. For example, different language codes correspond to different encryption algorithm selections, encryption parameter settings, etc. A mapping relationship table can be constructed based on the language codes. After obtaining the codes of the first language and the second language, the corresponding encryption algorithm is searched in the table. For example, a code combination corresponds to encrypting using the AES (Advanced Encryption Standard) algorithm, and specific encryption mode details are determined according to the characteristics of the language. For a language with a more complex grammar structure, a segmented encryption mode may be selected, where the voice data is segmented according to semantic units or grammar structures and then encrypted separately; for a language with a higher lexical richness, a local encryption method based on specific word frequencies is determined, etc., so as to generate an encryption mode suitable for the voice data.

[0108] By combining language codes to generate the encryption mode, it is possible to fully consider the characteristic differences of different languages in multilingual interactions, make the encryption method more suitable for the actual situation of voice data, avoid using a general encryption method, effectively improve the security and effectiveness of voice data encryption, provide a basic guarantee for the subsequent secure transmission of voice data, and ensure the information security during the voice interaction process.

[0109] The communication key is a key element for encrypting and decrypting voice data. Relying solely on a fixed and general key is prone to security risks and cannot be closely associated with the specific voice interaction content. Generating the communication key based on the semantic analysis results can make the key have the characteristic of being related to the current voice interaction semantics, further improving the security of encryption, and at the same time ensuring the key fits the specific voice data content. Extract various valuable semantic information from the semantic analysis results, such as key topic words, semantic structure features, semantic roles, etc., and then convert this semantic information into key materials through a specific algorithm or model. This makes the encryption operation closely connected to the specific voice interaction semantics. Even if an attacker obtains some encrypted voice data, it is very difficult to crack the original voice content without understanding the corresponding semantics and the generated key. At the same time, it also avoids problems such as the traditional fixed key being easily cracked and the security being reduced due to repeated use, provides a reliable and personalized key guarantee for the encryption of voice data, and enhances the performance of the entire voice interaction system in terms of communication security.

[0110] In the previous two steps, the encryption mode is determined and the communication key is generated respectively. Then, the two are combined to perform the actual encryption operation on the reply voice data, converting the original plaintext voice data into ciphertext form, and then sending the encrypted voice data to the user terminal. During the encryption process, according to the established encryption mode (such as specific methods like segmented encryption, local encryption, etc.), the communication key is used to encrypt each part of the voice data through the corresponding encryption algorithm (such as AES in symmetric encryption algorithms or asymmetric encryption algorithms, etc., depending on the specific encryption mode selection), transforming the characteristic information of the voice data (such as the values corresponding to the audio waveform data, etc.) into ciphertext content that cannot be directly recognized and understood, ensuring that even if intercepted during transmission, third parties are difficult to obtain the real information therein.

[0111] After the encryption is completed, through means such as network communication protocols, the encrypted voice data is sent according to information such as the address of the target terminal, enabling it to be accurately transmitted to the user's terminal device, such as sending it to the user's smartphone, smart speaker and other terminals, waiting for the corresponding decryption operation at the receiving end to restore the voice data, and realizing secure voice interaction communication.

[0112] Through the actual encryption operation and correct feedback transmission, the privacy and security of the voice data during the transmission process from the system end to the user terminal are effectively protected, preventing security problems such as information leakage and tampering, enabling users to use the multimodal voice interaction service with confidence, and enhancing the reliability and user trust of the entire system.

[0113] In one embodiment, generating the voice data encryption mode based on the codes of the first language and the second language includes:

[0114] Based on the language code mapping table, map the codes of the first language and the second language to unique first digital identifiers and second digital identifiers respectively;

[0115] Calculate the first difference and the second difference between the first digital identifier, the second digital identifier and a preset encryption base value respectively;

[0116] Encrypt the sum value of the first difference and the second difference based on a preset encryption algorithm to obtain an encrypted characteristic value;

[0117] Perform binary decomposition on the encrypted characteristic value, and analyze it in combination with the frequency and duration of the reply voice data, and determine the encryption segmentation method and local encryption area through the dynamic programming algorithm.

[0118] In this embodiment, in a scenario where multimodal voice interaction involves multiple languages, in order to generate an appropriate voice data encryption mode based on language information, first, a standardized and unique way is needed to represent different languages. The language code mapping table is a pre-set set of corresponding relationships constructed for this purpose. It corresponds the code of each language (which can be in the form of a specific encoding, symbol, etc.) to a unique digital identifier. The principle behind this is that numbers are more convenient for mathematical operations and following established algorithmic logics in subsequent calculations and processing. Through the mapping operation, the language codes with abstract semantics are transformed into digital forms that can be easily processed by a computer, laying the foundation for generating the encryption mode subsequently.

[0119] This mapping method enables different languages to have a unified and standardized digital representation form in the subsequent encryption mode generation process, facilitating various mathematical operations and logical judgments, and avoiding the inconvenience brought by directly processing the relatively complex language codes that are difficult to directly participate in calculations. It is the starting step for generating the encryption mode based on language information. Subsequent steps rely on these accurate digital identifiers to carry out, ensuring the orderliness and operability of the encryption mode generation process.

[0120] The preset encryption base value is a fixed value pre-set by the system. By calculating the differences between the digital identifiers corresponding to the first language and the second language and this encryption base value, a variable factor related to the language can be introduced. The principle lies in that the digital identifiers obtained by mapping different languages are different, and the differences obtained by subtracting the encryption base value are also different. These differences can carry the characteristic information of the language itself and serve as unique parameters in the subsequent encryption process to affect the generation of the encryption mode, enabling the encryption mode to be associated with specific languages rather than adopting a general and fixed way, increasing the pertinence and diversity of encryption.

[0121] By calculating the differences, the language differences represented by the language codes are cleverly transformed into quantifiable numerical differences that can participate in subsequent encryption operations, enabling the finally generated encryption mode to vary according to different language combinations, better adapting to the personalized requirements of voice data encryption in a multilingual environment, avoiding potential security vulnerabilities of a single fixed encryption mode, and enhancing the security and flexibility of encryption.

[0122] After obtaining the first difference and the second difference related to language, sum them up and perform encryption processing using a preset encryption algorithm. The purpose is to further integrate the language-related information and transform it into a confidential encrypted eigenvalue through the encryption operation. The encryption algorithm used here can be a common symmetric encryption algorithm (such as AES, etc.) or an asymmetric encryption algorithm (such as RSA, etc.), specifically depending on the overall security design requirements of the system. The encryption process is to take the sum of the differences as the input data according to the rules of the selected encryption algorithm, and through the key in the encryption algorithm (this key is pre-configured by the system for the encryption operation at this stage, different from the communication key used for encrypting voice data subsequently) and specific encryption transformation operations, convert it into an encrypted eigenvalue in ciphertext form. The above-mentioned encrypted eigenvalue not only integrates the information related to the first language and the second language (reflected by the differences), but also has confidentiality after encryption processing. It becomes an important intermediate data for further determining the specific encryption method of voice data (such as segmentation method, local encryption area, etc.), and plays a key role in connecting the preceding with the following in the entire encryption mode generation chain.

[0123] On the one hand, the difference information related to language is encrypted and integrated, enabling the language characteristics to participate in the subsequent process in a secure and encrypted form; on the other hand, it provides a unified and encrypted key data basis for finally determining the specific details of voice data encryption, ensuring that the process from generating intermediate data based on language to determining the actual encryption method is secure and logically coherent, and guaranteeing the integrity and security of the encryption mode generation.

[0124] Performing binary decomposition on the encrypted eigenvalue is to convert it into a binary form that is more convenient for the computer to analyze and process bit by bit, so as to extract information that can be used to guide the specific encryption operation of voice data. And analyzing in combination with the frequency and duration information of the replied voice data itself takes into account the characteristics of voice data in terms of acoustic properties. Different frequency bands and duration intervals may have different requirements and sensitivities for encryption. The dynamic programming algorithm plays a role here. Based on the principle of an optimal substructure, it decomposes the problem into multiple sub-problems, records the optimal solutions of each sub-problem, and finally finds the optimal solution of the entire problem. In this scenario, it is to use the dynamic programming algorithm to comprehensively consider the language-related encryption information contained in the encrypted eigenvalue after binary decomposition and the frequency and duration characteristics of voice data to determine how to perform segmented encryption on voice data and delimit the local encryption area, so as to achieve the purpose of being able to make full use of language-related information to ensure the pertinence of encryption and combine the characteristics of the voice itself to achieve efficient and reasonable encryption.

[0125] For example, a combination of several bits of the encrypted eigenvalue after binary decomposition can indicate that the voice data is segmented according to the frequency, and different encryption strengths or encryption algorithm details are used for the low-frequency band, the middle-frequency band, and the high-frequency band respectively. At the same time, according to the voice duration and other information in the encrypted eigenvalue, the dynamic programming algorithm is used to calculate which specific time periods in the voice data are more suitable for local encryption. For example, the time periods containing key semantic information are encrypted with emphasis, so as to achieve an encryption mode that is more in line with the actual situation of the voice data and the language characteristics.

[0126] The above steps finally determine the specific implementation method of voice data encryption. By combining various factors, especially fully considering the language-related information and the acoustic characteristics of the voice data itself, a personalized encryption scheme is generated, avoiding the problems of over-encryption (affecting efficiency) or insufficient encryption (posing security risks) that may be caused by the general encryption method to the voice data. The encryption mode can maximize the security and transmission efficiency of the voice data in the multi-language and multi-modal voice interaction environment, which is the key closing link in the process of generating the voice data encryption mode based on the language code number and is directly related to the actual effect of voice data encryption.

[0127] In one embodiment, generating a communication key based on the semantic analysis result includes:

[0128] Extracting the lexical feature vector, syntactic feature vector, and semantic dependency relationship feature vector from the semantic analysis result;

[0129] Constructing a graph structure based on each feature vector; wherein, the lexical feature vector is used as a node, and the syntactic feature vector and the semantic dependency relationship feature vector are used to determine the weights of the edges;

[0130] Performing a depth-first traversal on the graph structure, and mapping the obtained node sequence to a numerical sequence;

[0131] Generating a polynomial curve using the numerical sequence, generating a key seed according to the coefficients and control point information of the polynomial curve; performing hashing on the key seed using a hash function to obtain a hash value;

[0132] Taking the hash value as an input, and generating a communication key through a key stream generation algorithm based on a linear feedback shift register.

[0133] In this embodiment, the semantic analysis result contains rich language information. To generate a communication key that is closely related to the specific semantics and has sufficient randomness and security, it is necessary to extract features from different perspectives. The lexical feature vector aims to capture the characteristics at the lexical level of the text, such as the part of speech, word frequency, semantic category, etc. of the words. Through a specific word vector representation method, each word is transformed into a vector representation, and these vectors together form the lexical feature vector, which can reflect the characteristics of the text in terms of word usage. The syntactic feature vector focuses on the syntactic structure of the text, such as the syntactic structure type of the sentence (subject-predicate-object, subject-linking verb-predicative, etc.), the arrangement order of syntactic components, the complexity of syntactic rules, etc. Through syntactic analysis tools and corresponding coding methods, the syntactic-related information is quantified into vector form to reflect the features at the syntactic level of the text. The semantic dependency relationship feature vector is mainly used to depict the dependency relationship between different semantic components in the text. For example, which words are the executors of actions and which are the recipients. By using semantic dependency analysis technology to determine the relationship between each semantic unit and transforming it into a vector representation, the deep semantic structure features of the text can be mined.

[0134] The above three feature vectors quantitatively represent the semantic analysis result from three key dimensions of vocabulary, syntax, and semantics, providing rich and multi-faceted basic materials for subsequent construction of the graph structure and generation of the communication key. This enables the finally generated communication key to fully integrate the language characteristics of the text, avoiding problems such as weak correlation and insufficient randomness that may occur when generating the key based on a single dimension. It is an important starting step in the entire communication key generation process for the conversion from semantics to computable features.

[0135] Constructing the extracted different feature vectors into a graph structure is to present the internal relationship between various language elements in the semantic analysis result in an intuitive and convenient way for subsequent processing. Data structures such as the adjacency matrix or adjacency list in graph theory can be used to implement the construction of the graph. Taking the lexical feature vector as a node means that each word has a corresponding node entity in the graph, while the syntactic feature vector and the semantic dependency relationship feature vector are used to determine the weight of the edge. This is because syntactic rules determine the degree of connection tightness between words in the sentence structure, and semantic dependency relationships reflect the strength of the semantic association between words. By comprehensively considering these two feature vectors to set the weight of the edge, the association situation between different words based on syntax and semantics can be accurately reflected. For example, if two words are in a direct subject-predicate relationship grammatically and have a close semantic dependency (such as "bird" and "sing"), then the weight of the edge between them will be relatively high, and vice versa. The graph structure constructed in this way is like a semantic relationship network, clearly showing the internal language logical structure of the text.

[0136] By constructing a graph structure, the feature vectors of different dimensions extracted previously are effectively integrated into an organic whole, which more intuitively shows the language logical relationships in the semantic analysis results, provides a clear structural basis for further mining and utilizing these semantic association information to generate communication keys through operations such as graph traversal in the follow-up, and is a key step in transforming scattered feature vectors into semantic relationship carriers available for key generation.

[0137] Depth-first traversal is a common traversal algorithm in graph theory. Selecting it to traverse the constructed graph structure is to access the nodes in the graph (i.e., the words corresponding to the lexical feature vectors) in a specific order, so as to mine the information of the semantic association order reflected by the graph structure. During the traversal process, the nodes are recorded in the order of access to form a node sequence, and this sequence carries the order information of the text semantics unfolding according to the logical relationships in the graph. Then this node sequence is mapped to a numerical sequence. Usually, a pre-set coding rule can be adopted. For example, a unique digital code is assigned to each word corresponding to the lexical feature vector (the coding value can be determined based on the vocabulary order, word frequency ranking, etc.), and the corresponding codes are combined according to the order of the words in the node sequence to obtain the numerical sequence. The purpose of doing this is to further transform the semantic order information reflected by the graph structure into a numerical form convenient for subsequent mathematical processing, in preparation for key generation.

[0138] Through depth-first traversal and mapping operations, the semantic features integrating syntax and semantic associations are orderly extracted from the graph structure and transformed into a numerical sequence, realizing the transformation from the semantic logical structure to a numerical representation available for mathematical operations, enabling the subsequent use of these numerical information to generate keys based on mathematical methods, and is an important intermediate step in further quantifying and computabilizing semantic information to serve communication key generation.

[0139] Generating a polynomial curve using a numerical sequence is to further explore the hidden rules and features in the numerical sequence by leveraging the characteristics of mathematical functions. Based on the numerical values in the numerical sequence as the sampling points of the polynomial curve, the coefficients of the polynomial are determined through mathematical methods such as fitting. For example, curve fitting algorithms such as the least squares method are used to find a polynomial function that can pass through these sampling points well. The coefficients of this polynomial contain the key feature information of the numerical sequence. At the same time, some key control points on the polynomial curve (such as extreme points, inflection points, etc.) also carry unique information. By synthesizing this coefficient and control point information to generate a key seed, the key seed is closely related to the numerical sequence transformed from the previous semantic analysis result, that is, related to the semantics of the text. Then, a hash function is used to perform a hash operation on the key seed. The hash function has characteristics such as one-wayness and collision resistance. By performing a hash operation on the key seed, it is converted into a hash value with a fixed length. This hash value has high randomness and confidentiality, while retaining the relevance to the original key seed (and thus to the semantic analysis result), providing a relatively secure and appropriate intermediate value for the final generation of the communication key.

[0140] A linear feedback shift register (LFSR) is a stream cipher generation mechanism commonly used to generate a key stream. It is based on the principle of a shift register. Through feedback logic, shift and feedback operations are performed on the initial value in the register (i.e., the hash value obtained previously), continuously generating a new bit sequence, and this bit sequence is the key stream. In practical applications, a sequence with an appropriate length and format is extracted from this key stream according to certain rules (such as grouping, etc.) as the final communication key. The principle is that the LFSR can utilize the small differences in the initial value to generate a key stream with good randomness and a long period. Here, the hash value obtained through multiple previous steps of semantic-related processing is used as the initial value, so that the generated communication key not only has semantic relevance based on semantics but also inherits the security characteristics such as randomness and periodicity brought by the LFSR, meeting the requirements for the security and confidentiality of the key in the communication process.

[0141] By using the key stream generation algorithm of the LFSR, the hash value obtained through complex processing previously is converted into a finally usable communication key, so that the generated communication key integrates the semantic features of the semantic analysis result and good security characteristics, can effectively encrypt voice data, etc., and ensure the communication security in the multi-modal voice interaction process. It is the key final step in the entire communication key generation technical solution.

[0142] In one embodiment, generating a communication key based on the semantic analysis result includes:

[0143] Encoding the semantic analysis result to obtain encoded characters;

[0144] Insert the encoded characters into each node in the graph structure one by one in sequence, and assign values to the edges between adjacent nodes according to the relationship between the characters on the adjacent nodes to obtain a character graph structure;

[0145] Find all the edges with different assigned values as target edges; connect the midpoints of each target edge in sequence to obtain multiple target connections;

[0146] Find the nodes that satisfy a preset relationship with each target connection in the character graph structure as target nodes;

[0147] Combine the characters on the target nodes in sequence as a key seed, and perform hashing on the key seed using a hash function to obtain a hash value;

[0148] Use the hash value as input and generate a communication key through a key stream generation algorithm based on a linear feedback shift register.

[0149] In this embodiment, the semantic analysis result is usually natural language content presented in text form. To facilitate subsequent processing in specific mathematical models such as graph structures and follow the established algorithm logic, it needs to be converted into a standardized and easy-to-operate encoding form, that is, encoded characters. This encoding process can be implemented based on various methods. For example, common character encoding standards (such as ASCII code, Unicode, etc.), or encoding rules designed according to specific application scenarios, convert each character, word, or semantic unit in the text into a corresponding encoded value according to a certain mapping relationship. The character forms corresponding to these encoded values are encoded characters. The advantage of this is to convert the relatively abstract and complex language information in the semantic analysis result into a discrete symbol form that can be easily recognized and processed by a computer, providing basic materials for subsequent operations such as constructing a graph structure.

[0150] Inserting encoded characters into the nodes of the graph structure is to organize and present the information in the semantic analysis results in a structured way. The graph structure can well reflect the mutual relationships between elements. Here, the semantic elements represented by the encoded characters are carried by the nodes. Assigning values to the edges between adjacent nodes according to the relationships between the characters on the adjacent nodes is to further explore and quantify the degree of association between these semantic elements. This relationship can be measured from multiple perspectives, such as the similarity of characters (which can be judged by the numerical difference of character encodings, the syntactic and semantic relevance of characters, etc.), the sequential relationship of characters in the text, etc. According to these relationships, corresponding numerical values are assigned to the edges through preset assignment rules (for example, the higher the similarity, the greater the assigned value, and the assigned value for adjacent sequences is a specific value, etc.). The character graph structure constructed in this way is like a semantic network that integrates semantic encoding information and semantic element association information. Each node carries the encoded semantic part, and the assignment of the edges reflects their closeness or other association characteristics.

[0151] In this embodiment, the encoded semantic information is effectively incorporated into the graph structure, and the relationships between semantic elements are quantified through the assignment of edges, making the entire structure not only contain the semantics itself but also reflect its internal logical connection, providing a clear structural basis for subsequent operations such as finding edges, nodes with specific properties, etc. It is a key link for further exploring the semantics that can be used to generate communication key information, enabling subsequent steps to perform targeted processing based on this character graph structure that integrates information from multiple aspects.

[0152] Furthermore, finding edges with non-identical assignments as target edges is to screen out those edges with unique association characteristics from the character graph structure. These edges represent the distinctive connection relationships between semantic elements and are of great value for generating communication keys with randomness and uniqueness. Connecting the midpoints of these target edges to obtain the target connection line is to further explore the potential connections between these unique edges. By connecting the lines, a new structural clue is constructed. This clue can connect some elements in the graph in a new way based on the unique semantic associations represented by the target edges, and it is possible to discover more rules or characteristics that are more conducive to key generation.

[0153] For example, in a constructed character graph structure, if there are several edges with assignments of 3, 4, 5, etc. (and no edges with the same assignment), select the edges with different assignments such as the edge with an assignment of 3 and the edge with an assignment of 4 as target edges. Then find the midpoint of each target edge (the midpoint position can be determined by geometric methods such as calculating the average of the coordinates of the two end nodes of the edge). Connect these midpoints in sequence (such as in the natural order of the edges in the graph or the order sorted based on the assignment size, etc.), and multiple target connections are obtained. These connections form a new geometric structure in the graph, containing new information based on unique edge relationships.

[0154] By screening target edges and constructing target connections, unique associated information is extracted from the numerous edges and relationships in the character graph structure. This screened and reorganized information becomes a key basis with high distinctiveness and good randomness when generating communication keys later, which helps improve the tight correlation between the communication key and the semantic analysis result, as well as the security and unpredictability of the key itself. It is an important step in further mining key information for key generation based on the character graph structure.

[0155] The above preset relationship is a matching rule between nodes and target connections preset according to the requirements of generating communication keys and the semantic association characteristics carried by the character graph structure. By finding the nodes that satisfy this preset relationship with each target connection, the aim is to focus on the semantic elements (carried by nodes) that have a close connection with the previously mined unique connections. For example, the preset relationship can be set such that the vertical distance from a node to a target connection is within a certain range (measuring the relevance from a geometric perspective), or there is a certain specific dependency relationship in semantics between the encoded characters represented by the node and the encoded characters of the two end nodes of the target connection, etc. The target nodes selected according to such rules concentrate the semantic information closely related to the unique semantic associations mined through the target edges and target connections before, and become the key source for generating the key seed.

[0156] When the target connections are determined, for each node in the character graph structure, according to the preset distance relationship rule, calculate the distance from the node to each target connection. If the distance from a certain node to a certain target connection meets the set range requirements, then this node is determined as a target node. By performing such screening and judgment on all nodes in the entire character graph structure, all target nodes that meet the requirements are found.

[0157] By accurately locating the target nodes according to the preset relationship, semantic elements closely related to the previously mined unique association structure are extracted from the character graph structure, enabling the subsequently generated key seeds to maximally integrate the key semantic information screened and mined through multiple rounds in the character graph structure, ensuring a deep association between the communication key and the semantic analysis result, further enhancing the pertinence and security of the key, and providing core semantic materials for generating high-quality communication keys.

[0158] Combining the characters on the target nodes in sequence to form key seeds is because these target nodes carry the key semantic information mined and screened through multiple steps from the semantic analysis result. Combining them forms a character sequence closely related to semantics. Using this as the key seed gives the seed the uniqueness and relevance derived from the semantic analysis result. Then, a hash function is used to perform hash processing on the key seed. The characteristics of the hash function, such as one-wayness and collision resistance, can convert the key seed into a hash value with a fixed length, high randomness, and confidentiality. This not only retains the relevance to semantics (because the hash value is derived from the key seed generated based on semantics) but also endows it with security through the hash function, providing a suitable intermediate value for generating the final communication key.

[0159] The key stream generation algorithm of the linear feedback shift register (LFSR) is based on the principle of the shift register. Through specific feedback logic, shift and feedback operations are performed on the initial value (the above-mentioned hash value) in the register to continuously generate a new bit sequence, which is the key stream. In practical applications, a sequence with a suitable length and format is extracted from this key stream according to certain rules (such as grouping, etc.) as the final communication key. The principle is that the LFSR can utilize the small differences in the initial value to generate a key stream with good randomness and a long period. Here, using the hash value obtained through multiple previous semantic-related processes as the initial value makes the generated communication key not only have the relevance based on semantics but also inherit the security characteristics such as randomness and periodicity brought by the LFSR, meeting the requirements for the security and confidentiality of the key in the communication process.

[0160] By using the key stream generation algorithm of the LFSR, the hash value obtained through complex processing previously is converted into the finally available communication key, making the generated communication key integrate the semantic features of the semantic analysis result and good security characteristics, and being able to effectively encrypt voice data, etc., to ensure the communication security in the multi-modal voice interaction process. It is the key final step of the entire communication key generation technical solution.

[0161] Refer to Figure 2 , an embodiment of the present invention also provides a multi-modal voice interaction device based on a large model, including:

[0162] An identification module, configured to perform speech recognition on the speech data in the first language input by the user to obtain the first language text;

[0163] An analysis module, configured to perform language translation on the first language text, translate it into the second language text; perform semantic analysis on the second language text to obtain a semantic analysis result;

[0164] A construction module, configured to perform multi-modal knowledge graph construction on the semantic analysis result based on a preset large model to obtain an enhanced semantic understanding result; wherein, the multi-modal knowledge graph integrates multi-modal information related to the second language text;

[0165] A generation module, configured to generate a dialogue strategy for the enhanced semantic understanding result;

[0166] A synthesis module, configured to perform speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language.

[0167] In this embodiment, for the specific implementation of each module in the above device embodiment, please refer to that described in the above method embodiment, and details will not be elaborated here.

[0168] Refer to Figure 3 , in an embodiment of the present invention, an electronic device is further provided. The internal structure of the electronic device may be as Figure 3 shown. The electronic device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store the corresponding data in this embodiment. The network interface of the electronic device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.

[0169] Those skilled in the art can understand that Figure 3 the structure shown in

[0170] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the electronic device to which the solution of the present invention is applied.

[0171] In summary, the multi-modal voice interaction method, electronic device, and storage medium based on a large model provided in the embodiments of the present invention include: performing speech recognition on the speech data of the first language input by the user to obtain the first language text; performing language translation on the first language text to translate it into the second language text; performing semantic analysis on the second language text to obtain a semantic analysis result; constructing a multi-modal knowledge graph based on a preset large model for the semantic analysis result to obtain an enhanced semantic understanding result, where the multi-modal knowledge graph integrates multi-modal information related to the second language text; generating a dialogue strategy for the enhanced semantic understanding result; and performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language. In the present invention, speech recognition and language translation are performed on the speech data of the first language input by the user to obtain the second language text, then semantic analysis is performed to obtain a semantic analysis result, and a multi-modal knowledge graph is constructed, and finally reply speech data is generated. It realizes interaction in multiple languages and integrates various modal information at the same time.

[0172] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0173] It should be noted that in this document, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.

[0174] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A multimodal voice interaction method based on a large model, characterized in that: The following steps are involved: Performing speech recognition on the speech data in the first language input by the user to obtain a text in the first language; performing language translation on the first language text into a second language text; Performing semantic analysis on the second language text to obtain a semantic analysis result; Based on a preset large model, a multimodal knowledge graph is constructed for the semantic analysis result to obtain an enhanced semantic understanding result; wherein the multimodal knowledge graph integrates multimodal information related to the second language text; Generating a dialogue strategy for enhancing the semantic understanding result; Performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language; A multimodal knowledge graph is constructed for the semantic analysis results based on a preset large model to obtain enhanced semantic understanding results, including: Encoding the semantic analysis result based on the text encoder in the large model to obtain a text vector; For multimodal information related to the second language text, using a corresponding image encoder and a video encoder to extract and encode features of the image data and the video data, respectively, to obtain an image feature vector and a video feature vector; The image feature vector and the video feature vector are fused into the text vector to construct an initial knowledge graph; wherein the text vector is a core node, and the fused multimodal features are used as peripheral nodes and edges to form an initial knowledge graph; The initial knowledge graph is optimized based on a semantic relationship model to obtain an enhanced semantic understanding result; wherein, the semantic relationship model is based on a multimodal corpus for unsupervised learning, can identify and supplement the semantic relationships in the knowledge graph, and improve the structure and content of the knowledge graph by continuously iteratively adjusting the connection relationship and semantic annotations between nodes.

2. The multimodal voice interaction method based on a large model according to claim 1, characterized in that: The dialogue strategy for generating the enhanced semantic understanding result includes: Based on a preset intent classification model, the enhanced semantic understanding result is divided into a plurality of intent categories; wherein the intent classification model is obtained by pre-training; For each intent category, a corresponding dialogue strategy template is matched from a pre-stored dialogue strategy template library; the dialogue strategy template library is pre-built based on multi-language and multi-scenario dialogue data; By using natural language generation technology, the matched dialogue strategy template is filled and adjusted according to the specific content in the enhanced semantic understanding result to generate a personalized dialogue strategy.

3. The multimodal voice interaction method based on a large model according to claim 1, characterized in that: The image feature vector and the video feature vector are merged into the text vector, comprising: Through the cross-modal attention mechanism, the text vector interacts with the image feature vector and the video feature vector to calculate the association weights between different modalities; wherein the cross-modal attention mechanism is based on a multi-head attention architecture and captures the association information between different modalities by calculating multiple attention heads in parallel; According to the calculated association weight, the image feature vector and the video feature vector are fused into the text vector.

4. The multimodal voice interaction method based on a large model according to claim 1, characterized in that: After performing speech synthesis based on the dialogue strategy to generate reply speech data expressed in the first language, the method includes: Generate a voice data encryption pattern based on the codes of the first language and the second language; generating a communication key based on the semantic analysis result; Based on the voice data encryption mode, the reply voice data is encrypted using the communication key and then fed back to the user's terminal.

5. The multimodal voice interaction method based on a large model according to claim 4 is characterized in that: The method of generating a voice data encryption mode based on the codes of the first language and the second language includes: Based on the language code mapping table, the codes of the first language and the second language are mapped into a unique first digital identifier and a unique second digital identifier respectively; Calculating a first difference and a second difference between the first digital identifier, the second digital identifier and a preset encryption base value respectively; Encrypting the sum of the first difference and the second difference based on a preset encryption algorithm to obtain an encrypted feature value; The encrypted feature value is binary decomposed, and analyzed in combination with the frequency and duration of the reply voice data, and the encrypted segmentation method and local encryption area are determined through a dynamic programming algorithm.

6. The multimodal voice interaction method based on a large model according to claim 4 is characterized in that: Generating a communication key based on the semantic analysis result includes: Extracting vocabulary feature vectors, grammatical feature vectors and semantic dependency feature vectors from the semantic analysis results; A graph structure is constructed based on each feature vector, wherein the vocabulary feature vector is used as a node, and the grammatical feature vector and the semantic dependency feature vector are used to determine the weight of the edge; Performing a depth-first traversal on the graph structure, and mapping the node sequence obtained by the traversal into a numerical sequence; Generate a polynomial curve using a numerical sequence, generate a key seed according to the coefficients of the polynomial curve and the control point information; hash the key seed using a hash function to obtain a hash value; The hash value is used as input to generate the communication key through a key stream generation algorithm based on a linear feedback shift register.

7. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Vehicle man-machine voice interaction method and system and vehicle

    CN118968992A

  • Multi-modal document structured processing and knowledge extraction method based on large language model

    CN119227794A