Voice recognition-based interaction method, apparatus and device, and storage medium

Through the interaction method based on voice recognition, users' voices are recognized and personalized reply are generated, the shortcomings of the existing system in terms of intelligence and emotional interaction are solved, and highly personalized and emotionally delicate human-computer interaction is achieved, which improves user satisfaction.

CN120199252APending Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510446634.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing human-computer interaction system has shortcomings in terms of intelligence level and emotional interaction capabilities, which is difficult to meet the diversified needs of the silver-haired group, the children group and the female group. The emotional recognition and feedback mechanism is immature, affecting service effects and user satisfaction.

Method used

Using an interaction method based on speech recognition, the trained recognition model converts user speech into voice text, and recognizes the speaker's identity and user emotions. The language model generates personalized reply text, and realizes voice interaction through the text-to-speech model.

Benefits of technology

It realizes highly personalized, intelligent and emotionally delicate human-computer interaction, improves user satisfaction and user stickiness, and can more accurately identify and respond to user emotions and needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199252A_ABST
    Figure CN120199252A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses an interaction method and device based on voice recognition, equipment and a storage medium, and the method comprises the steps: obtaining collected user voice; a recognition model is adopted to convert the user voice into voice characters, and the identity of a speaker and the emotion of the user are obtained through recognition; obtaining the above-mentioned reply text, converting the voice text, the speaker identity and the user emotion into a prompt word text by adopting a language model, and generating a target reply text according to the above-mentioned reply text and the prompt word text; and converting the target dialogue text into target dialogue voice by adopting a text-to-voice model, and controlling a loudspeaker to play the target dialogue voice. The method can be applied to business management program systems of financial science and technology, medical treatment and the like, the problem of double limitation of intelligence and emotional interaction faced by existing human-computer interaction is solved, and the industry service quality and the user stickiness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology and is applied to online business processing scenarios such as fintech and digital healthcare. In particular, it relates to an interaction method, device, equipment, and storage medium based on speech recognition. Background Art

[0002] Existing interaction systems such as chatbots, companion robots, and intelligent customer services generally have deficiencies in the two dimensions of intelligence level and emotional interaction ability, making it difficult to meet the diverse needs of the elderly, children, women, and other groups in specific economic fields.

[0003] In terms of intelligence level, existing interaction systems generally face the dilemma of insufficient learning ability and are difficult to dynamically optimize their own performance by deeply mining continuous interaction data with users, resulting in the difficulty of synchronously improving the accuracy and efficiency of service responses as user needs evolve. In addition, the system lacks the ability of deep personalized customization and cannot adjust refined interaction strategies according to user identity characteristics (such as age, gender, financial investment preferences, health status, etc.) and preferences, making the service experience tend to be homogenized and difficult to establish long-term user stickiness.

[0004] On the emotional interaction level, the emotional recognition and feedback mechanism of existing interaction systems is still immature and can only handle basic emotional expressions, making it difficult to accurately respond to the complexity and dynamics of human emotions. In scenarios such as financial consulting and medical consultations that highly rely on emotional resonance and trust building, this limitation of emotional interaction is particularly prominent, not only weakening the warmth and humanistic care of the service, but also possibly greatly reducing the service effect due to emotional misjudgment, thereby affecting user satisfaction and loyalty.

[0005] Therefore, developing an interaction system that can break through the dual limitations of existing intelligence and emotional interaction, and achieve highly personalized, intelligent, and emotionally delicate interaction has become the key for various industries to improve service quality and enhance user stickiness. Summary of the Invention

[0006] The present invention provides an interaction method, device, equipment, and storage medium based on speech recognition to solve the problem of dual limitations of intelligence and emotional interaction faced by existing human-computer interaction.

[0007] In a first aspect, an interaction method based on speech recognition is provided, including:

[0008] Obtain the collected user speech;

[0009] Use the trained recognition model to convert the user speech into speech text, and identify the speaker identity and user emotion of the user speech;

[0010] Obtain the foregoing conversation text, and use the trained language model to convert the speech text, the speaker identity, and the user emotion into a prompt text, and generate a target conversation text according to the foregoing conversation text and the prompt text;

[0011] Use a text-to-speech model to convert the target conversation text into a target conversation voice, and control the speaker to play the target conversation voice.

[0012] In a second aspect, there is provided an interactive device based on speech recognition, including:

[0013] A speech acquisition module for acquiring the collected user speech;

[0014] A speech recognition module for using the trained recognition model to convert the user speech into speech text, and identifying the speaker identity and the user emotion of the user speech;

[0015] A text generation module for obtaining the foregoing conversation text, using the trained language model to convert the speech text, the speaker identity, and the user emotion into a prompt text, and generating a target conversation text according to the foregoing conversation text and the prompt text;

[0016] A speech conversion module for using a text-to-speech model to convert the target conversation text into a target conversation voice, and controlling the speaker to play the target conversation voice.

[0017] In a third aspect, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the foregoing interactive method based on speech recognition are implemented.

[0018] In a fourth aspect, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the foregoing interactive method based on speech recognition are implemented.

[0019] In the solution implemented by the above voice recognition-based interaction method, device, computer device, and storage medium, the collected user voice can be converted into voice text by using the trained recognition model, and the speaker identity and user emotion of the user voice can be recognized, which can more accurately recognize and understand the user input, help provide personalized services according to the speaker identity subsequently, and can more accurately understand the user's emotional intention based on the user emotion, facilitating corresponding emotional responses; by using the trained voice model to convert the voice text, speaker identity, and user emotion into a prompt text, and combining the foregoing conversation text and the prompt text to generate a target conversation text, the change of user interaction data can be deeply mined according to the context information provided by the foregoing conversation text, and a target conversation text that meets the evolution of user needs, user identity characteristics, and emotional response requirements can be generated by combining the voice text, speaker identity, and user emotion; by using the text-to-speech model to convert the target conversation text into a target conversation voice and controlling the speaker to play the target conversation voice, the voice interaction between the system and the user is realized, enhancing the convenience and comfort of the interaction. Based on this solution, through the combination of the foregoing interaction data, user identity characteristics, and emotion response requirements, the voice interaction of the input user voice is realized, achieving a highly personalized, intelligent, and emotionally delicate human-computer interaction, and improving user satisfaction and user stickiness. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0021] Figure 1 is an exemplary system architecture diagram of a voice recognition-based interaction method in an embodiment of the present invention;

[0022] Figure 2 is a flowchart of a voice recognition-based interaction method in an embodiment of the present invention;

[0023] Figure 3 is Figure 1 a specific implementation flowchart of step S60 in

[0024] Figure 4 is Figure 1 a specific implementation flowchart before step S60 in

[0025] Figure 5 is a structural diagram of a voice recognition-based interaction device in an embodiment of the present invention;

[0026] Figure 6 It is a schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0028] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0029] In order to enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings.

[0030] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0031] The user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0032] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and the like.

[0033] The server 103 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal device 101.

[0034] It should be noted that the voice recognition-based interaction method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the voice recognition-based interaction device is generally set in the server / terminal device.

[0035] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0036] Continue to refer to Figure 2 as shown in Figure 2 FIG. 18 is a schematic flowchart of a voice recognition-based interaction method provided by an embodiment of the present invention, including the following steps:

[0037] S20: Obtain the collected user voice;

[0038] The voice recognition-based interaction method provided by the present invention can be applied to the intelligent interaction engine in various application scenarios. The intelligent interaction engine is usually implemented by a server. The server is connected to the client through a network. The server can obtain the collected user voice through the client. Among them, the client can include, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.

[0039] Specifically, obtaining the collected user voice can be, for example, to capture the user voice in real time through a microphone or other audio collection devices; or to obtain the user voice pre-saved in the storage unit of the client.

[0040] Optionally, after obtaining the user voice, the user voice can also be pre-processed, including noise reduction, gain control, echo cancellation, etc., to improve the voice quality.

[0041] Exemplarily, in the financial industry, a user sends a voice consultation request, such as "I want to know about the latest fund products", to an intelligent voice customer service system through a device such as a smart phone or a smart speaker. The intelligent voice customer service system collects the user voice through a microphone and performs pre-processing such as noise reduction and echo cancellation to ensure the voice quality.

[0042] S40: Using the trained recognition model, convert the user voice into voice text, and identify the speaker identity and user emotion of the user voice;

[0043] In this embodiment, a trained recognition model is obtained. The recognition model is composed of three major functional module sets, namely, a voice-to-text conversion module, an identity recognition module, and an emotion recognition module. Among them, the voice-to-text conversion module is used to convert the user voice into voice text, the identity recognition module is used to identify the speaker identity in the user voice, and the emotion recognition module is used to identify the user emotion in the user voice.

[0044] Specifically, the obtained user voice is input into the trained recognition model. Through this recognition model, the user voice is converted into voice text, and the speaker identity and user emotion of the user voice are identified. For example, it can be that through the voice-to-text conversion module in the recognition model, the user voice is converted into voice text; through the identity recognition module in the recognition model, the speaker identity in the user voice is identified according to the user voice; through the emotion recognition module in the recognition model, the user emotion in the user voice is identified according to the user voice.

[0045] In some embodiments, the above voice-to-text conversion module is implemented by a trained speech recognition model, the identity recognition module is implemented by a trained voiceprint recognition model, and the emotion recognition module is implemented by a trained emotion recognition model. That is, the recognition model is constructed by a speech recognition model, a voiceprint recognition model, and an emotion recognition model. Step S40, that is, using the trained recognition model to convert the user voice into voice text and identify the speaker identity and user emotion of the user voice, may include the following steps:

[0046] Using the trained speech recognition model, extract the target speech features of the user voice and convert the target speech features into voice text;

[0047] Specifically, using the trained speech recognition model, target speech features are extracted from the user's speech, such as Mel Frequency Cepstral Coefficients (MFCC), Log-Filterbank Energies, etc., and the extracted target speech features are converted into speech text. Among them, the speech recognition model can adopt the Automatic Speech Recognition (ASR) model, or other models such as Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Transformer, or models based on the attention mechanism.

[0048] Using the trained voiceprint recognition model, the target speech features are matched with the pre-registered historical speech features to identify the speaker's identity;

[0049] Specifically, the target speech features can also include voiceprint features. Using the trained Voiceprint Recognition (VPR) model, the extracted target speech features (including voiceprint features) are subjected to embedding learning to obtain target embedding vectors; historical embedding vectors of the historical speaker are extracted from the pre-registered historical speech features; the target embedding vectors are matched with the historical embedding vectors, and the speaker's identity is obtained according to the matching result.

[0050] Using the trained emotion recognition model, emotion classification is performed according to the target speech features to obtain the user's emotion.

[0051] Specifically, the target speech features also include features such as pitch, energy, and speech rate. Using the trained emotion recognition model, the probability scores of each emotion category are calculated according to the extracted target speech features, and the emotion category with the highest probability score is selected as the user's emotion. Among them, the emotion recognition model can be constructed using a Deep Neural Network (DNN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) and its variants (such as LSTM, GRU), etc.

[0052] In this embodiment, by training the speech recognition model, voiceprint recognition model, and emotion recognition model separately and optimizing them for their respective tasks, the three tasks of speech recognition, voiceprint recognition, and emotion recognition can be completed simultaneously, significantly improving the recognition efficiency. Moreover, each model focuses on its specific task, so that relevant information can be extracted and analyzed more precisely.

[0053] S60: Obtain the aforementioned conversation text, use the trained language model to convert the speech text, the speaker's identity, and the user's emotion into a prompt text, and generate a target conversation text according to the aforementioned conversation text and the prompt text;

[0054] In this embodiment, the foregoing conversation text is obtained. The foregoing conversation text refers to the text record of the conversation content that has occurred between the user and the system before the current conversation turn. The foregoing conversation text is the key basis for the system to understand the current conversation context, identify the conversation topic, and the user's intention. Taking the recognized speech text, the speaker identity, and the user's emotion as inputs, and using the trained language model, a prompt text is generated according to the speech text, the speaker identity, and the user's emotion. This prompt text is used to summarize the user's current request or intention and guide the language model to generate a conversation text that conforms to the context and the user's intention. Through natural language processing (NLP) technology, context analysis is performed on the foregoing conversation text to understand the historical background, topic events, and key entity relationships of the current conversation, etc.; specific requests, attention preferences, and emotion needs of the user are extracted from the prompt text. For example, the specific request of the user is determined according to the recognized speech text, the attention preference is determined according to the speaker identity, and the emotion need is determined according to the user's emotion. According to the above analysis results of the foregoing conversation text and the prompt text, a target conversation text is generated.

[0055] In some embodiments, step S60, that is, using the trained language model to convert the speech text, the speaker identity, and the user's emotion into a prompt text, may include the following steps:

[0056] Using the trained language model, perform word segmentation on the speech text to obtain a number of word segmentation results;

[0057] Specifically, using the trained language model, clean the speech text to remove noise information such as irrelevant characters, punctuation marks, or numbers, etc., to obtain pure text content; use a rule-based method, a statistical method, or a deep learning-based method to perform word segmentation on this pure text content to obtain a number of word segmentation results.

[0058] Obtain a preset identity list and a numerical representation mapping table, match the identity list and the numerical representation mapping table with the speaker identity, and obtain the first numerical representation corresponding to the speaker identity;

[0059] Specifically, obtain a preset identity list and a numerical representation mapping table. This identity list and numerical representation mapping table contains a number of historical speaker identities and the numerical representations corresponding to these historical speaker identities. Match the identity list and the numerical representation mapping table with the recognized speaker identity, that is, match each of the several historical speaker identities in the mapping table with the recognized speaker identity one by one. When there is a match, obtain the numerical representation corresponding to the matched historical speaker identity as the first numerical representation of the recognized speaker identity.

[0060] Obtain a preset emotion classification and numerical representation mapping table, match the emotion classification and numerical representation mapping table with the user emotion, and obtain the second numerical representation corresponding to the user emotion;

[0061] Specifically, obtain a preset emotion classification and numerical representation mapping table, which includes several emotion classifications and the corresponding numerical representations of these emotion classifications. Match the emotion classification and numerical representation mapping table with the recognized user emotion, that is, match each of the several emotion classifications in the mapping table with the recognized user emotion one by one. When a match occurs, obtain the numerical representation corresponding to the matched emotion classification as the second numerical representation of the recognized user emotion.

[0062] Obtain a preset prompt word template, and correspondingly fill the word segmentation result, the first numerical representation, and the second numerical representation into the prompt word template to generate a prompt word text.

[0063] Specifically, obtain a preset prompt word template, which contains placeholders for filling the word segmentation result, the first numerical representation corresponding to the speaker identity, and the second numerical representation corresponding to the user emotion. Correspondingly fill the word segmentation result, the first numerical representation, and the second numerical representation into the prompt word template, that is, fill the word segmentation result, the first numerical representation (representing the speaker identity), and the second numerical representation (representing the user emotion) into the placeholders of the prompt word template to generate a prompt word text.

[0064] In some embodiments, referring to Figure 3 , step S60, that is, generating a target reply text according to the foregoing reply text and the prompt word text, may include the following steps S601 - S604:

[0065] S601: Adopt a trained language model to capture the context information in the foregoing reply text based on the attention layer;

[0066] Specifically, adopt a trained language model to analyze the semantic relationships and dependencies in the foregoing reply text by using the attention layer inside the model to obtain the context information in the foregoing reply text. Among them, the attention layer can automatically focus on different parts of the input text, dynamically allocate attention weights, and thus effectively capture the context information in the text.

[0067] S602: Map the context information and the word segmentation result in the prompt word text into a pre - constructed word vector space, and gradually predict the next word to generate a reply text sequence;

[0068] Specifically, obtain the word segmentation results in the prompt text, map the context information of the foregoing conversation text and the word segmentation results in the prompt text into a pre-constructed word vector space, capture the semantic relationships between each word through this word vector space, and convert the words into high-dimensional vectors; use the mapped high-dimensional vectors, combined with the sequence generation ability of the language model, to gradually predict the next word and generate a conversation text sequence.

[0069] S603: Match the language style preference strategy according to the first numerical representation in the prompt text, and match the emotion coping strategy according to the second numerical representation in the prompt text;

[0070] Specifically, define corresponding language style preference strategies in advance according to different speaker identities and store them in the language style preference strategy library. According to the first numerical representation (representing the speaker identity) in the prompt text, perform a match in the pre-defined language style preference strategy library to obtain the language style preference strategy corresponding to this first data representation. This language style preference strategy can include different language styles such as formal, informal, professional, and cordial.

[0071] Specifically, define corresponding emotion coping strategies in advance according to different emotion classifications and store them in the emotion coping strategy library. According to the second numerical representation (representing the user's emotion) in the prompt text, perform a match in the pre-defined emotion coping strategy library to obtain the emotion coping strategy corresponding to this second data representation. This emotion coping strategy can include different emotion coping methods such as comfort, concern, encouragement, humor, and seriousness.

[0072] S604: Perform rhetorical adjustment on the conversation text sequence according to the language style preference strategy and the emotion coping strategy, and output the target conversation text.

[0073] Specifically, combine the language style preference strategy and the emotion coping strategy to perform rhetorical adjustment on the generated conversation text sequence to obtain the target conversation text. Among them, the rhetorical adjustment can include word selection, sentence structure adjustment, choice of tone strength, addition of polite expressions, and assistance of modal particles, etc., so that the generated target conversation text is more natural, fluent and in line with the context.

[0074] Exemplarily, when the generated conversation text sequence is "It may get colder tomorrow. Keep warm to prevent catching a cold", if the speaker's identity is female and her language style preference strategy is kind and her emotion coping strategy is concerned, the generated target conversation text should be adjusted to "It may get colder tomorrow. Wear more clothes to keep warm and don't catch a cold"; if the speaker's identity is an elderly person and her language style preference strategy is formal and her emotion coping strategy is serious, the generated target conversation text should be adjusted to "There is a high possibility of temperature drop tomorrow. It is recommended to reduce going out. If you go out, please pay attention to keeping warm to prevent catching a cold".

[0075] In this embodiment, by capturing the context information in the aforementioned conversation text through the attention layer, the continuity and context of the conversation can be better understood, which helps to generate a target conversation text that is coherent and logically reasonable with the previous conversation; by mapping the context information and the word segmentation results in the prompt text into the word vector space and gradually predicting the next word, words and phrases can be flexibly combined according to the given context and prompt to generate diverse conversation texts; by rhetorically adjusting the conversation text sequence according to the language style preference strategy and emotion coping strategy, the generated target conversation text is not only accurate in content, but also more appropriate in expression and emotional color, enhancing the expressiveness and appeal of the conversation text. Through the organic combination of context understanding, personalized conversation generation and strategic rhetorical adjustment, personalized, coherent and expressive target conversation texts are generated, significantly improving the user's interaction experience and satisfaction.

[0076] In some embodiments, it may also include the process of training the speech model. Specifically, in step S20, that is, before acquiring the collected user speech, the following steps may also be included:

[0077] Acquire pre-collected historical interaction information, where the historical interaction information includes historical conversation text, historical speech text, historical speaker identity, historical user emotion and standard conversation text;

[0078] Specifically, acquire pre-collected historical interaction information, which is a historical data record generated during the interaction between the historical user and the system, and includes historical conversation text, historical speech text, historical speaker identity, historical user emotion and standard conversation text. Among them, the historical conversation text refers to the text record of the conversation content that has occurred between the historical user and the system before the conversation round represented by the historical speech text. The standard conversation text is the ideal reply content marked by the system or manually as a reference target.

[0079] Adopt the initial language model and capture the historical context information in the historical conversation text based on the attention layer;

[0080] Specifically, an initial language model is adopted to analyze the semantic relationships and dependencies in the historical conversation text using the attention layer inside the model, so as to obtain the historical context information in the historical conversation text.

[0081] Perform word segmentation on the historical speech text to obtain a historical word segmentation result, map the historical word segmentation result and the historical context information to the initial word vector space, and gradually predict the next word to generate a conversation text sequence for training;

[0082] Specifically, clean the historical speech text to remove noise information such as irrelevant characters, punctuation marks or numbers; use a rule-based method, a statistical method or a deep learning-based method to perform word segmentation on the cleaned historical speech text to obtain several historical word segmentation results. Map the historical word segmentation result and the historical context information into the initial word vector space, capture the semantic relationships between each word through the initial word vector space, and convert the words into high-dimensional vectors; use the mapped high-dimensional vectors, combined with the sequence generation ability of the language model, to gradually predict the next word and generate a conversation text sequence for training.

[0083] Assign a third numerical representation to the historical speaker identity, randomly obtain a training language style preference strategy from a preset style preference strategy library, and construct a first matching relationship according to the historical speaker identity, the third numerical representation and the training language style preference strategy;

[0084] Specifically, assign a third numerical representation to the historical speaker identity. For example, extract the identification information of the historical speaker identity (such as user ID, role label, etc.), convert the identification information into a numerical form that can be processed by the model. For example, use an embedding vector to map the identification information into a low-dimensional vector space to obtain the third numerical representation. This third numerical representation facilitates the model to capture the semantic information of the historical speaker identity. Randomly obtain a training language style preference strategy from a pre-constructed style preference strategy library, and associate the historical speaker identity, the third numerical representation with the training language style preference strategy to construct a first matching relationship.

[0085] Assign a fourth numerical representation to the historical user emotion, randomly obtain a training emotion coping strategy from a preset emotion coping strategy library, and construct a second matching relationship according to the historical user emotion, the fourth numerical representation and the training emotion coping strategy;

[0086] Specifically, assign a fourth numerical representation to the historical user emotion. For example, extract the emotion tags of the historical user emotion (such as anger, happiness, calmness, etc.), and convert the emotion tags into a numerical form that can be processed by the model. For example, use one-hot encoding to encode the emotion tags to obtain the fourth numerical representation. This fourth numerical representation facilitates the model to capture the semantics and intensity of the historical user emotion. Randomly obtain the emotion coping strategies for training from the pre-constructed emotion coping strategy library, and associate the historical user emotion, the fourth numerical representation with the emotion coping strategies for training to construct the second matching relationship.

[0087] According to the training language style preference strategy and the training emotion coping strategy, perform rhetorical adjustment on the training conversation text sequence, and output the training conversation text;

[0088] Specifically, combine the training language style preference strategy and the training emotion coping strategy to perform rhetorical adjustment on the training conversation text sequence, and output the training conversation text. Among them, the rhetorical adjustment can include word selection, sentence structure adjustment, selection of tone strength, addition of polite expressions, assistance of modal particles, etc., so that the generated target conversation text is more natural, fluent and in line with the context.

[0089] Obtain a preset loss function, and calculate the loss value in combination with the training conversation text, the standard conversation text and the loss function;

[0090] Specifically, obtain a preset loss function, such as cross-entropy loss, mean square error, Euclidean distance, etc.; perform calculations in combination with the training conversation text, the standard conversation text and the loss function. For example, combine the Euclidean distance calculation formula to calculate the distance between the training conversation text and the standard conversation text as the loss value.

[0091] Update the model parameters of the initial language model according to the loss value;

[0092] Specifically, calculate the gradient according to the calculated loss value, and use an optimization algorithm (such as gradient descent, Adam, etc.) to perform backpropagation according to the gradient to update the model parameters of the initial speech model.

[0093] Perform iterative training on the language model after parameter update until the model converges. Save the first matching relationship and the second matching relationship after the model converges, and output the trained language model.

[0094] Specifically, the updated language model is iteratively trained multiple times until the model converges (i.e., the loss value no longer changes significantly). The first matching relationship after the model converges (i.e., the association relationship between the historical speaker identity, the third numerical representation, and the language style preference strategy used for training) and the second matching relationship (i.e., the association relationship between the historical user emotion, the fourth numerical representation, and the emotion coping strategy used for training) are saved, and the trained language model is output.

[0095] Optionally, the first matching relationship can be used to construct an identity list and a numerical representation mapping table, and to define the language style preference strategies corresponding to different speaker identities. The second matching relationship can be used to construct an emotion classification and numerical representation mapping table, and to define the emotion coping strategies corresponding to different emotion classifications.

[0096] In some embodiments, to further enhance the accuracy of user authentication, dual authentication of identity can also be performed by combining the user's voice and user image. Specifically, referring to Figure 4 , in step S60, that is, before converting the speech text, the speaker identity, and the user emotion into a prompt text using the trained language model, the following steps S501 - S503 can also be included:

[0097] S501: Obtain the collected user image;

[0098] S502: Use the trained visual model to identify the user object identity based on the user image;

[0099] S503: Match the speaker identity and the user object identity. If they match, perform the step of converting the speech text, the speaker identity, and the user emotion into a prompt text using the trained language model.

[0100] Specifically, the user image is collected in real time through a camera or other image acquisition device. The trained visual model (such as YOLO, Faster R - CNN, etc.) is used to extract the image features of the user image, and the user object identity is identified based on the image features. The speaker identity recognized by the language model is matched with the user object identity. If the speaker identity matches the user object identity, it indicates that the user's voiceprint is consistent with the user's image. At this time, the step of converting the speech text, the speaker identity, and the user emotion into a prompt text using the trained language model is performed.

[0101] In this embodiment, by obtaining a user image and using a trained vision model to identify the identity of the user object, a visual authentication method is added. The combination of this visual authentication and the speaker identity (i.e., obtained through voiceprint recognition) improves the accuracy and reliability of user authentication. By matching the speaker identity with the user object identity, it is confirmed that the user's voiceprint is consistent with the user's image. This dual authentication method further enhances the credibility of authentication and effectively prevents fraud and impersonation in special business scenarios such as finance and healthcare.

[0102] In some embodiments, to further improve the user experience and satisfaction, voice recognition and visual recognition can also be combined to optimize the human-computer interaction strategy. Specifically, in step S60, that is, after generating the target conversation text according to the foregoing conversation text and the prompt word text, the following steps may further be included:

[0103] Using the trained vision model, identify the user's action based on the user image;

[0104] Specifically, for the collected user image, use the trained vision model to extract action features from the user image, and identify the user's action based on the action features, such as actions like gestures, postures, nodding or shaking the head.

[0105] Generate an interaction response strategy based on the user's action and the target conversation text;

[0106] Specifically, generate an interaction response strategy based on the user's action and the target conversation text. For example, extract key information in the target conversation text, such as system responses, emotional responses, etc., and establish an association between the user's action and the key information of the target conversation text. For example, when the user's action is a victory gesture, the system response in the target conversation text is "You did a great job" and its emotional response is "positive", the association relationship can be victory gesture - great - positive emotion. Generate an interaction response strategy based on this association relationship, and this interaction response strategy can include an interaction expression response strategy and an interaction action response strategy, etc.

[0107] Control the corresponding interaction response component to make an interaction response action according to the interaction response strategy.

[0108] Specifically, control the corresponding interaction component to make an interaction response action according to the generated interaction response strategy. The interaction component may include but is not limited to screen display, animation playback, robot actions, robot expression display, lighting control, etc.

[0109] In this embodiment, by generating an interactive response strategy based on the user's actions and the target conversation text, not only the user's actions themselves are considered, but also the key information in the conversation text, such as system responses and emotion handling, is incorporated. Thus, the correlation between the user's actions and text information is established, making the system's interactive responses more in line with the user's needs and context. By accurately identifying the user's actions, intelligently generating interactive response strategies, and flexibly controlling interactive components, a more personalized, natural, and smooth interactive experience can be provided for the user. This high degree of interactivity and responsiveness helps enhance the user's sense of participation and satisfaction, thereby improving the overall usage effect of the system.

[0110] S80: Use a text-to-speech model to convert the target conversation text into target conversation speech, and control the speaker to play the target conversation speech.

[0111] Specifically, use a text-to-speech model (such as a Text To Speech, TTS model) to perform preprocessing on the target conversation text, such as removing special characters and converting numbers or letters. Then convert the preprocessed target conversation text into target conversation speech, which exists in the form of a voice waveform file (such as WAV, MP3, etc.) or streaming audio data. Transmit the voice waveform file or streaming audio data to the speaker to trigger the speaker to play the target conversation speech.

[0112] Exemplarily, in the customer service area of a financial institution, a cute pet-shaped robot designed specifically for the financial scenario is deployed. This robot not only has a cute appearance but also has powerful speech recognition and interaction capabilities, aiming to provide a more cordial and personalized service experience for customers.

[0113] When a customer says to the robot in the customer service area of a financial institution: "I want to check my account balance." The robot's microphone quickly captures the customer's voice information and, through the trained recognition model, converts the voice into text: "I want to check my account balance." At the same time, the model also identifies the speaker's identity (such as regular customer Mr. Zhang) and the user's emotion (such as calm and expectant).

[0114] Next, the robot obtains the aforementioned conversation text, such as "I want to purchase XX financial product", and uses the trained language model, combined with the speaker's identity and the user's emotion, to generate a prompt text: "Mr. Zhang, with a relatively calm and expectant emotion, wants to check account balance information." Then, based on the aforementioned conversation text and the prompt text, the robot generates a target conversation text: "Mr. Zhang, your account balance is still XX yuan. Do you need to know more about XX financial product?"

[0115] During this process, the robot also uses the trained visual model to capture the customer's actions in real time. When the customer makes a thumbs-up gesture, the robot quickly recognizes it and decides to adopt an interactive expression response strategy. The robot's eye screen displays a heart expression, and the speaker plays the voice: "Thank you for your recognition, we will continue to work hard to provide you with better service."

[0116] For example, in the outpatient hall of the smart hospital, a cute pet-like robot designed for medical scenarios is busy interacting with patients. An elderly patient walked up to the robot and said in a slightly trembling voice: "I came for a follow-up check today, but I don't know where to go." The robot's microphone keenly captured the patient's voice and quickly converted the voice into text through the trained recognition model: "I came for a follow-up check today, but I don't know where to go." At the same time, the model also recognized the patient's speaker identity (such as Aunt Li who often comes for a follow-up check) and the user's emotions (such as some confusion and anxiety).

[0117] The robot then obtains the aforementioned reply text, which can be empty this time. Then, the robot uses the trained language model, combined with the speaker's identity and user emotions, to generate a prompt text: "Aunt Li needs a reexamination, but she doesn't know the location of the department for the reexamination. She is confused and anxious." Then, the robot generates a target reply text based on the prompt text: "Aunt Li, which department do you want to go to for a reexamination?" Aunt Li replied "Internal Medicine Clinic". The robot generates the prompt text again: "Aunt Li needs to go to the Internal Medicine Clinic for a reexamination, but she doesn't know the location of the department. She is confused and anxious." Then, based on the aforementioned reply text and the prompt text "Aunt Li needs to go to the Internal Medicine Clinic for a reexamination, but she doesn't know the location of the department. She is confused and anxious", the robot generates the target reply text: "Aunt Li, don't worry, you should go to the third floor of the Internal Medicine Clinic for a reexamination. I'll take you there."

[0118] Then, the robot used its built-in navigation system to guide Aunt Li to the third floor of the Internal Medicine Clinic. While walking, Aunt Li accidentally tripped and her body swayed slightly. The robot quickly sensed this through gesture recognition technology, and immediately stretched out its "arm" to support Aunt Li, saying with concern: "Aunt Li, be careful, I will help you walk." At the same time, the robot's eye screen showed a worried expression, making Aunt Li feel cared for and taken care of.

[0119] Through this interactive method, the cute pet-like robot not only provides patients with accurate and timely guidance services, but also makes patients feel cared for and warm through rich expressions and action responses. This interactive method based on voice recognition plays an important role in smart medical scenarios and improves patients' medical experience and satisfaction.

[0120] The interactive method based on speech recognition provided by the embodiments of the present invention uses a trained recognition model to convert the collected user speech into speech text, and identifies the speaker identity and user emotion of the user speech, which can more accurately recognize and understand the user input, helps to subsequently provide personalized services according to the speaker identity, and can more accurately understand the user's emotional intention based on the user emotion, facilitating corresponding emotional responses; by using the trained speech model to convert the speech text, speaker identity and user emotion into prompt text, and generating a target reply text by combining the foregoing reply text and prompt text, it can deeply mine the changes in user interaction data according to the context information provided by the foregoing reply text, and combine the speech text, speaker identity and user emotion to generate a target reply text that meets the evolution of user needs, user identity characteristics and emotional response requirements; by using a text-to-speech model to convert the target reply text into a target reply voice and controlling the speaker to play the target reply voice, it realizes the voice interaction between the system and the user, enhancing the convenience and comfort of the interaction. Based on this solution, by combining the foregoing interaction data, user identity characteristics, and emotion response requirements to perform voice interaction on the input user speech, it realizes a highly personalized, intelligent, and emotionally delicate human-computer interaction, improving user satisfaction and user stickiness.

[0121] It should be emphasized that to further ensure the privacy and security of the above user speech, the foregoing reply text, and the target reply text and other information, the above user speech, the foregoing reply text, and the target reply text and other information can also be stored in a node of a blockchain.

[0122] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods, and each data block contains information about a batch of network transactions, used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0123] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems.

[0124] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0126] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit and can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be completed at the same moment, but can be executed at different moments. Their execution order does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0127] In one embodiment, an interaction device based on speech recognition is provided. The interaction device based on speech recognition corresponds one-to-one with the interaction method based on speech recognition in the above embodiment. As Figure 5 shown, the interaction device based on speech recognition includes a speech acquisition module 201, a speech recognition module 202, a text generation module 203, and a speech conversion module 204. The detailed description of each functional module is as follows:

[0128] The speech acquisition module 201 is used to acquire the collected user speech;

[0129] The speech recognition module 202 is used to convert the user speech into speech text by using a trained recognition model, and identify the speaker identity and user emotion of the user speech;

[0130] The text generation module 203 is configured to obtain the foregoing conversation text, and use the trained language model to convert the speech text, the speaker identity, and the user emotion into a prompt text, and generate a target conversation text according to the foregoing conversation text and the prompt text;

[0131] The speech conversion module 204 is configured to use a text-to-speech model to convert the target conversation text into a target conversation speech, and control a speaker to play the target conversation speech.

[0132] In one embodiment, the text generation module 203 is specifically configured to:

[0133] Use the trained language model to perform word segmentation on the speech text to obtain a plurality of word segmentation results;

[0134] Obtain a preset identity list and a numerical representation mapping table, match the identity list and the numerical representation mapping table with the speaker identity, and obtain a first numerical representation corresponding to the speaker identity;

[0135] Obtain a preset emotion classification and numerical representation mapping table, match the emotion classification and numerical representation mapping table with the user emotion, and obtain a second numerical representation corresponding to the user emotion;

[0136] Obtain a preset prompt template, and correspondingly fill the word segmentation results, the first numerical representation, and the second numerical representation into the prompt template to generate a prompt text.

[0137] In one embodiment, the text generation module 203 is specifically configured to:

[0138] Use the trained language model to capture the context information in the foregoing conversation text based on the attention layer;

[0139] Map the context information and the word segmentation results in the prompt text into a pre-constructed word vector space, and gradually predict the next word to generate a conversation text sequence;

[0140] Match a language style preference strategy according to the first numerical representation in the prompt text, and match an emotion coping strategy according to the second numerical representation in the prompt text;

[0141] According to the language style preference strategy and the emotion coping strategy, perform rhetorical adjustment on the conversation text sequence, and output a target conversation text.

[0142] In one embodiment, the voice recognition-based interaction device is further configured to:

[0143] Obtain pre-collected historical interaction information, where the historical interaction information includes historical conversation text, historical speech text, historical speaker identity, historical user emotion, and standard conversation text;

[0144] Adopt an initial language model to capture historical context information in the historical conversation text based on the attention layer;

[0145] Perform word segmentation on the historical speech text to obtain a historical word segmentation result, map the historical word segmentation result and the historical context information to the initial word vector space, and gradually predict the next word to generate a conversation text sequence for training;

[0146] Assign a third numerical representation to the historical speaker identity, randomly obtain a training language style preference strategy from a preset style preference strategy library, and construct a first matching relationship based on the historical speaker identity, the third numerical representation, and the training language style preference strategy;

[0147] Assign a fourth numerical representation to the historical user emotion, randomly obtain a training emotion coping strategy from a preset emotion coping strategy library, and construct a second matching relationship based on the historical user emotion, the fourth numerical representation, and the training emotion coping strategy;

[0148] Adjust the rhetoric of the training conversation text sequence according to the training language style preference strategy and the training emotion coping strategy, and output the training conversation text;

[0149] Obtain a preset loss function, and calculate a loss value in combination with the training conversation text, the standard conversation text, and the loss function;

[0150] Update the model parameters of the initial language model according to the loss value;

[0151] Perform iterative training on the language model after parameter update until the model converges, save the first matching relationship and the second matching relationship after the model converges, and output the trained language model.

[0152] In one embodiment, the interaction device based on speech recognition is further configured to:

[0153] Obtain the collected user image;

[0154] Adopt the trained visual model to identify the user object identity according to the user image;

[0155] Match the speaker identity and the user object identity. If they match, perform the step of using the trained language model to convert the speech text, the speaker identity, and the user emotion into a prompt word text.

[0156] In one embodiment, the voice recognition-based interaction device is further configured to:

[0157] Adopt the trained visual model to recognize the user's actions based on the user image;

[0158] Generate an interaction response strategy according to the user's actions and the target reply text;

[0159] Control the corresponding interaction response component to make an interaction response action according to the interaction response strategy.

[0160] In one embodiment, the recognition model is constructed by a voice recognition model, a voiceprint recognition model, and an emotion recognition model. The voice recognition module 202 is specifically configured to:

[0161] Adopt the trained voice recognition model to extract the target voice features of the user's voice and convert the target voice features into voice text;

[0162] Adopt the trained voiceprint recognition model to match the target voice features with the pre-registered historical voice features to identify the speaker's identity;

[0163] Adopt the trained emotion recognition model to perform emotion classification according to the target voice features to obtain the user's emotion.

[0164] The present invention provides a voice recognition-based interaction device. By adopting the trained recognition model, the collected user voice is converted into voice text, and the speaker's identity and the user's emotion of the user voice are recognized, which can more accurately recognize and understand the user input, help to provide personalized services according to the speaker's identity subsequently, and can more accurately understand the user's emotional intention based on the user's emotion, facilitating corresponding emotional responses; by using the trained voice model to convert the voice text, the speaker's identity, and the user's emotion into a prompt text, combining the foregoing reply text and the prompt text to generate a target reply text, it can deeply mine the changes in user interaction data according to the context information provided by the foregoing reply text, and combine the voice text, the speaker's identity, and the user's emotion to generate a target reply text that meets the evolution of user needs, the user identity characteristics, and the emotional response requirements; by adopting a text-to-speech model, the target reply text is converted into a target reply voice, and the speaker is controlled to play the target reply voice to realize the voice interaction between the system and the user, enhancing the convenience and comfort of the interaction. Based on this solution, through the voice interaction of the input user voice by combining the foregoing interaction data, user identity characteristics, and emotion response requirements, highly personalized, intelligent, and emotionally delicate human-computer interaction is realized, improving user satisfaction and user stickiness.

[0165] For the specific limitations of the voice recognition-based interaction device, reference can be made to the limitations of the voice recognition-based interaction method in the above text, which will not be elaborated here. Each module in the above voice recognition-based interaction device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0166] To solve the above technical problems, an embodiment of the present application further provides a computer device. For details, please refer to Figure 6 , Figure 6 , which is the basic structural block diagram of the computer device in this embodiment.

[0167] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 6 with a memory 61, a processor 62, and a network interface 63 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0168] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, or other means.

[0169] The memory 61 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, FlashCard, etc. equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions of the voice recognition-based interaction method. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.

[0170] In some embodiments, the processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the voice recognition-based interaction method.

[0171] The network interface 63 may include a wireless network interface or a wired network interface, and the network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0172] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions, and the computer-readable instructions can be executed by at least one processor, so that the at least one processor executes the steps of the voice recognition-based interaction method as described above.

[0173] The computer device, computer-readable storage medium, and computer-readable instructions provided by the embodiments of the present application can, by executing a trained recognition model through a processor, convert the collected user voice into voice text, and identify the speaker identity and user emotion of the user voice, so as to more accurately recognize and understand the user input, which helps to subsequently provide personalized services according to the speaker identity, and can more accurately understand the user's emotional intention based on the user emotion, facilitating corresponding emotional responses; by using the trained voice model to convert the voice text, speaker identity, and user emotion into prompt text, and combining the foregoing conversation text and prompt text to generate a target conversation text, it can deeply mine the changes in user interaction data based on the context information provided by the foregoing conversation text, and combine the voice text, speaker identity, and user emotion to generate a target conversation text that meets the evolution of user needs, user identity characteristics, and emotional response requirements; by adopting a text-to-speech model, converting the target conversation text into a target conversation voice, and controlling the speaker to play the target conversation voice, it realizes the voice interaction between the system and the user, enhancing the convenience and comfort of the interaction. Based on this solution, through combining the foregoing interaction data, user identity characteristics, and emotion response requirements, the input user voice is subjected to voice interaction, realizing highly personalized, intelligent, and emotionally delicate human-computer interaction, and improving user satisfaction and user stickiness.

[0174] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0175] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.

[0176] In the embodiments of this application, the software tools or components that are not of our company are only introduced by way of example and do not represent actual use.

Claims

1. An interactive method based on speech recognition, characterized in that: include: Obtaining the collected user voice; Using the trained recognition model, the user's voice is converted into speech text, and the speaker identity and user emotion of the user's voice are recognized; Obtain the aforementioned reply text, use the trained language model to convert the voice text, the speaker identity and the user emotion into prompt word text, and generate a target reply text according to the aforementioned reply text and the prompt word text; The target reply text is converted into a target reply voice by using a text-to-speech model, and a speaker is controlled to play the target reply voice.

2. The method according to claim 1, characterized in that The method of using the trained language model to convert the voice text, the speaker identity and the user emotion into a prompt word text includes: Using the trained language model, the speech text is segmented to obtain a number of segmentation results; Obtaining a preset identity list and a numerical representation mapping table, matching the identity list and the numerical representation mapping table with the speaker identity, and obtaining a first numerical representation corresponding to the speaker identity; Obtaining a preset emotion classification and numerical representation mapping table, matching the emotion classification and numerical representation mapping table with the user emotion, and obtaining a second numerical representation corresponding to the user emotion; A preset prompt word template is obtained, and the word segmentation result, the first numerical representation, and the second numerical representation are correspondingly filled into the prompt word template to generate a prompt word text.

3. The method according to claim 1, characterized in that The step of generating a target conversation text according to the conversation text and the prompt word text includes: Using the trained language model to capture contextual information in the aforementioned reply text based on an attention layer; Mapping the context information and the word segmentation results in the prompt word text into a pre-constructed word vector space, gradually predicting the next word, and generating a reply text sequence; According to the first numerical value in the prompt word text, the language style preference strategy is matched, and according to the second numerical value in the prompt word text, the emotional coping strategy is matched; According to the language style preference strategy and the emotional coping strategy, the reply text sequence is rhetorically adjusted to output a target reply text.

4. The method according to claim 1, characterized in that Before acquiring the collected user voice, the method further includes: Acquire pre-collected historical interaction information, wherein the historical interaction information includes historical reply text, historical voice text, historical speaker identity, historical user emotion and standard reply text; An initial language model is adopted to capture historical context information in the historical response text based on an attention layer; Perform word segmentation processing on the historical voice text to obtain a historical word segmentation result, map the historical word segmentation result and the historical context information to an initial word vector space, gradually predict the next word, and generate a training response text sequence; Assigning a third numerical representation to the historical speaker identity, randomly acquiring a language style preference strategy for training from a preset style preference strategy library, and establishing a first matching relationship according to the historical speaker identity, the third numerical representation, and the language style preference strategy for training; Assigning a fourth numerical representation to the historical user emotion, randomly acquiring an emotion coping strategy for training from a preset emotion coping strategy library, and establishing a second matching relationship according to the historical user emotion, the fourth numerical representation, and the emotion coping strategy for training; According to the language style preference strategy for training and the emotional coping strategy for training, rhetorically adjusting the reply text sequence for training, and outputting the reply text for training; Obtaining a preset loss function, combining the training reply text, the standard reply text and the loss function, and calculating a loss value; Updating the model parameters of the initial language model according to the loss value; The language model after parameter update is iteratively trained until the model converges, the first matching relationship and the second matching relationship after the model converges are saved, and the trained language model is output.

5. The method according to claim 1, characterized in that Before converting the voice text, the speaker identity and the user emotion into prompt word text using the trained language model, the method further includes: Obtaining the collected user image; Using the trained visual model, the user object identity is obtained according to the user image recognition; The speaker identity and the user object identity are matched, and if they match, the step of using the trained language model to convert the voice text, the speaker identity and the user emotion into prompt word text is executed.

6. The method according to claim 5, characterized in that After generating the target conversation text according to the conversation text and the prompt word text, the method further includes: Using the trained visual model, obtaining user actions according to the user image recognition; Generate an interactive response strategy according to the user action and the target conversation text; According to the interactive response strategy, the corresponding interactive response component is controlled to perform an interactive response action.

7. The method according to claim 1, characterized in that The recognition model is constructed by a speech recognition model, a voiceprint recognition model and an emotion recognition model. The trained recognition model is used to convert the user's speech into speech text, and the speaker identity and user emotion of the user's speech are recognized, including: Using the trained speech recognition model, extracting target speech features of the user's speech, and converting the target speech features into speech text; Using the trained voiceprint recognition model, the target voice features are matched with pre-registered historical voice features to identify the speaker; The trained emotion recognition model is used to perform emotion classification according to the target speech features to obtain user emotions.

8. An interactive device based on speech recognition, characterized in that: include: A voice acquisition module, used to acquire the collected user voice; A speech recognition module, used to convert the user's speech into speech text using a trained recognition model, and to identify the speaker's identity and the user's emotion of the user's speech; A text generation module, used to obtain the aforementioned reply text, use the trained language model to convert the voice text, the speaker identity and the user emotion into a prompt word text, and generate a target reply text according to the aforementioned reply text and the prompt word text; The speech conversion module is used to convert the target reply text into a target reply speech by adopting a text-to-speech model, and control the speaker to play the target reply speech.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the interaction method based on speech recognition as claimed in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the interactive method based on speech recognition as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Online man-machine conversation method, system and device, electronic equipment, storage medium and program product

    CN120723892A

  • Personalized voice interaction method and related equipment

    CN121053983A

  • Virtual standardized patient image generation and dialogue method and system

    CN121545653A

  • A method and system for virtual standardized patient avatar generation and dialogue

    CN121545653B

  • Interactive data processing method and device, computer equipment and readable storage medium

    CN121561814A