Voice interaction processing method and system, interaction management terminal and storage medium

By using an interactive management terminal and a cloud-based text processing model, cross-device voice interaction and linkage were achieved, solving the problems of collaboration and emotional interaction between smart terminal devices and improving user experience and interaction stability.

CN121838779APending Publication Date: 2026-04-10SHENZHEN KONKA ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the voice interaction functions of smart terminal devices are usually independent and closed, and cannot be coordinated across devices. This results in users needing to initiate voice commands separately for different devices, which is cumbersome and lacks interaction stability and emotional interaction.

Method used

An interactive management terminal is provided, which acquires interactive voice by detecting wake words, analyzes the voice content using a cloud-based text processing model, determines the terminal wake-up flag information, and stores the voice data locally to ensure stable transmission, thereby realizing cross-device voice interaction and linkage.

Benefits of technology

It simplifies user voice interaction, improves inter-device collaboration and ease of interaction, enhances user experience, enables emotional interactive feedback, and ensures the stability of the interaction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838779A_ABST
    Figure CN121838779A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction processing method and system, an interaction management terminal and a storage medium, and relates to the technical field of intelligent interaction, the method is applied to the interaction management terminal, and the method comprises the following steps: in response to a detected interaction wake-up word, obtaining an interaction voice input by a target object; terminal arousing mark information is determined according to the interaction voice, and the terminal arousing mark information is used for indicating a target interaction terminal needing to be aroused; and if the terminal arousing mark information indicates at least one target interaction terminal, sending the interaction voice to the target interaction terminal so as to trigger the target interaction terminal to interact with the target object according to the interaction voice. Therefore, the operation process of voice interaction of the user can be simplified, the convenience of voice interaction is improved, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent interaction technology, and in particular to a voice interaction processing method, system, interaction management terminal and storage medium. Background Technology

[0002] With the development of science and technology, especially the rapid development of artificial intelligence, voice interaction has become one of the mainstream ways for smart terminals to interact with users and is widely used in smart homes, smart offices and other scenarios.

[0003] In related technologies, voice interaction functions are typically tied to terminal devices, forming an independent interactive loop. This means that each device can only process its own voice interaction data independently, and the voice interaction capabilities of different devices cannot be coordinated. This results in users needing to issue voice commands separately for different devices, making the process cumbersome and hindering seamless cross-device collaborative interaction, thus reducing the convenience of voice interaction.

[0004] Therefore, the relevant technologies still need to be improved and developed. Summary of the Invention

[0005] The main purpose of this application is to provide a voice interaction processing method, system, interaction management terminal and storage medium, which aims to solve the technical problems in related technologies where devices can only process their own voice interaction data independently, the voice interaction capabilities between different devices cannot be coordinated, users need to initiate voice commands separately for different devices, the operation is cumbersome, and it is not conducive to improving the convenience of voice interaction.

[0006] To achieve the above objectives, the first aspect of this application provides a voice interaction processing method, which is applied to an interactive management terminal, and the method includes: In response to the detection of an interactive wake word, the interactive voice input from the target object is obtained; Based on the above interactive voice, the terminal wake-up flag information is determined, wherein the above terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up. If the aforementioned terminal wakes up the flag information to instruct at least one target interactive terminal, the aforementioned interactive voice is sent to the aforementioned target interactive terminal to trigger the aforementioned target interactive terminal to interact with the aforementioned target object based on the aforementioned interactive voice.

[0007] Optionally, the determination of terminal wake-up flag information based on the aforementioned interactive voice includes: The above-mentioned interactive voice is processed into text to obtain the interactive text corresponding to the interactive voice. Based on the interactive text, the terminal wake-up flag information is determined by a preset text processing model.

[0008] Optionally, the above text processing model is deployed on a cloud processing platform; The above-mentioned interactive voice is processed into text to obtain the interactive text corresponding to the interactive voice. Based on the interactive text, the terminal wake-up flag information is determined through a preset text processing model, including: The aforementioned interactive voice is uploaded to the cloud processing platform to trigger the cloud processing platform to perform text conversion processing on the aforementioned interactive voice, obtain the interactive text corresponding to the aforementioned interactive voice, determine the processing result based on the aforementioned interactive text through the aforementioned text processing model, and return the aforementioned processing result to the aforementioned interactive management terminal. The above processing result includes terminal activation flag information indicating the target interactive terminal; or, the above processing result includes emotion indication information and response text matching the content of the above interactive text, as well as terminal activation flag information that does not indicate any target interactive terminal.

[0009] Optionally, the aforementioned emotion indication information is used to indicate at least one target expression in a preset expression library.

[0010] Optionally, the above method further includes: If the aforementioned terminal wake-up flag does not indicate any target interactive terminal, then the aforementioned target emoticon is displayed, and audio is output according to the aforementioned response text.

[0011] Optionally, the target interactive terminal mentioned above is a smart TV; Sending the aforementioned interactive voice to the aforementioned target interactive terminal to trigger the target interactive terminal to interact with the aforementioned target object based on the aforementioned interactive voice includes: Send a voice assistant activation signal to the aforementioned smart TV to trigger the smart TV to activate the TV voice assistant; The aforementioned interactive voice is sent to the aforementioned TV voice assistant to trigger the aforementioned TV voice assistant to interact with the aforementioned target object based on the aforementioned interactive voice.

[0012] Optionally, the above-mentioned response to detecting an interactive wake word, acquiring the interactive voice input by the target object, includes: In response to the detection of an interactive wake word, the recording function of the aforementioned interactive management terminal is activated to record the interactive voice input by the target object, generate an audio file, and store it locally. Sending the aforementioned interactive voice to the aforementioned TV voice assistant to trigger the TV voice assistant to interact with the aforementioned target object based on the aforementioned interactive voice includes: The aforementioned audio file is sent to the aforementioned TV voice assistant to trigger the TV voice assistant to perform voice interaction with the aforementioned target object based on the content of the aforementioned audio file.

[0013] A second aspect of this application provides a voice interaction processing system, wherein the system is applied to an interactive management terminal, and the system includes: The data acquisition module is used to acquire the interactive voice input by the target object in response to the detection of an interactive wake word; The data processing module is used to determine terminal wake-up flag information based on the above interactive voice, wherein the terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up. An interaction processing module is used to send the interactive voice to the target interactive terminal if the terminal wakes up the flag information indicating at least one target interactive terminal, so as to trigger the target interactive terminal to interact with the target object according to the interactive voice.

[0014] A third aspect of this application provides an interactive management terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements any of the steps of the aforementioned voice interaction processing method.

[0015] A fourth aspect of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described voice interaction processing methods.

[0016] As can be seen from the above, the present application provides a voice interaction processing method applied to an interactive management terminal. Specifically, in response to detecting an interactive wake-up word, the method acquires the interactive voice input by the target object; based on the interactive voice, it determines terminal wake-up flag information, wherein the terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up; if the terminal wake-up flag information indicates at least one target interactive terminal, the method sends the interactive voice to the target interactive terminal to trigger the target interactive terminal to interact with the target object based on the interactive voice.

[0017] In this way, when engaging in voice interaction, users do not need to issue voice commands individually for each terminal device. Instead, they only need to wake up the interaction management terminal and input the interactive voice. The interaction management terminal will automatically determine the terminal wake-up flag information, thereby identifying the target interactive terminal to be activated, and sending the interactive voice to the target interactive terminal to trigger the interaction between the target terminal and the target object. Users only need to issue interactive voice to the interaction management terminal to link devices together and realize the invocation of other terminal devices for voice interaction. This simplifies the user's voice interaction operation process, improves the convenience of voice interaction, and enhances the user experience. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a voice interaction processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the specific process of a voice interaction processing method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a voice data processing flow provided in an embodiment of this application; Figure 4 This is a schematic diagram of the constituent modules of a voice interaction processing system provided in an embodiment of this application; Figure 5 This is a block diagram illustrating the internal structure of an interactive management terminal provided in an embodiment of this application. Figure 6 This is a six-view diagram of an interactive management terminal provided in an embodiment of this application; Figure 7 This is a schematic diagram of an application scenario for an interactive management terminal provided in an embodiment of this application. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.

[0021] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0022] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0023] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0024] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to classification." Similarly, the phrases "if determined" or "if classified to [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once classified to [the described condition or event]," or "in response to classification to [the described condition or event]."

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0026] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0027] Currently, voice interaction is being used more and more widely, and can be applied to various smart terminals. For example, as a core terminal for home entertainment and information access, the practicality and intelligence level of the voice interaction function of smart TVs directly affect the user experience.

[0028] Currently, while some smart TVs on the market are equipped with basic voice assistant functions, these functions generally have significant limitations. On the one hand, the level of intelligence in voice interaction is insufficient; most voice assistants can only respond to simple operation requests within a preset command set and cannot accurately understand and flexibly process the user's complex voice content. On the other hand, the interaction methods are monotonous and lack emotional expression capabilities; they cannot recognize the user's emotions based on the user's voice content and provide corresponding emotional feedback, making it difficult to meet users' needs for humanized interaction.

[0029] Meanwhile, in existing technologies, voice interaction functions are mostly tied to terminal devices, forming independent interactive loops, and the voice interaction capabilities between different devices cannot be coordinated. For example, the voice processing capabilities of smart speakers are disconnected from the display and playback capabilities of smart TVs, requiring users to issue voice commands separately for different devices, which is cumbersome and cannot achieve seamless cross-device collaborative interaction. In addition, some voice interaction solutions rely on real-time network transmission; if the network fluctuates or is interrupted, voice data can be easily lost, affecting the stability and reliability of the interaction.

[0030] Therefore, existing technologies suffer from problems such as low intelligence, lack of emotional engagement, poor cross-device collaboration, and unstable data transmission in smart terminal voice interaction. To address at least one of these technical problems, this application proposes a voice interaction processing method applied to an interaction management terminal. Specifically, in response to detecting an interaction wake-up word, the method acquires the interactive voice input by the target object; based on the interactive voice, it determines terminal wake-up flag information, wherein the terminal wake-up flag information is used to indicate the target interactive terminal to be woken up; if the terminal wake-up flag information indicates at least one target interactive terminal, the method sends the interactive voice to the target interactive terminal to trigger the target interactive terminal to interact with the target object based on the interactive voice.

[0031] In this way, when engaging in voice interaction, users do not need to issue voice commands individually for each terminal device. Instead, they only need to wake up the interaction management terminal and input the interactive voice. The interaction management terminal will automatically determine the terminal wake-up flag information, thereby identifying the target interactive terminal to be activated, and sending the interactive voice to the target interactive terminal to trigger the interaction between the target terminal and the target object. Users only need to issue interactive voice to the interaction management terminal to link devices together and realize the invocation of other terminal devices for voice interaction. This simplifies the user's voice interaction operation process, improves the convenience of voice interaction, and enhances the user experience.

[0032] like Figure 1 As shown in the figure, this application provides a voice interaction processing method applied to an interactive management terminal. Specifically, the method includes the following steps: Step S100: In response to detecting an interactive wake word, acquire the interactive voice input by the target object; Step S200: Based on the above interactive voice, determine the terminal wake-up flag information, wherein the terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up. Step S300: If the terminal wake-up flag information instructs at least one target interactive terminal, the interactive voice is sent to the target interactive terminal to trigger the target interactive terminal to interact with the target object based on the interactive voice.

[0033] The target object mentioned above is the object that initiates voice interaction commands through interactive voice. It can be a user or a device with voice interaction initiation function. In this embodiment of the application, the target object is a user as an example for specific explanation, but it is not intended to be a specific limitation.

[0034] Specifically, the interactive management terminal is pre-set with corresponding wake-up words. The aforementioned voice interaction processing method is only triggered when the user inputs the wake-up word, thus reducing the terminal's power consumption. The terminal has a built-in wake-up word detection module, which uses a neural network-based wake-up word model (such as the WakeNet model) specifically designed for low-power embedded scenarios. This module can monitor surrounding voice signals in real time while in standby mode, with low power consumption. Specifically, the wake-up word detection module can perform audio monitoring based on the Voice Activity Detection (VAD) algorithm to determine the voice activity state of the current frame in real time. It should be noted that the VAD algorithm can be used for preprocessing before wake-up word detection or for post-processing after wake-up word detection to optimize the entire interaction process. The preset wake-up word can be a custom phrase such as "start voice interaction," and the target (i.e., the user) can activate the interactive management terminal's voice interaction function by uttering this wake-up word.

[0035] Once the wake-up word detection module detects the interactive wake-up word, it immediately triggers the interactive management terminal to start recording, capturing the user's interactive voice input through the terminal's microphone array. The microphone array enables far-field voice acquisition, and combined with acoustic front-end (AFE) processing algorithms and acoustic echo cancellation (AEC) algorithms, it effectively removes interference from environmental noise and the user's own playback sound, ensuring clear voice signals and guaranteeing speech recognition even when the user is playing music or there is music or noise interference in the environment. Simultaneously, the interactive management terminal encodes the captured voice signal, generating an audio file, which is then stored in local storage to prevent data loss due to network issues. For example, Adaptive Differential Pulse Code Modulation (ADPCM) coding technology can be used to achieve efficient storage of voice data at a lower bit rate.

[0036] Specifically, the aforementioned terminal wake-up flag information is used to indicate the target interactive terminal to be woken up from a set of preset interactive terminals. It should be noted that the terminal wake-up flag information can indicate whether to wake up one target interactive terminal, multiple target interactive terminals, or not to wake up any target interactive terminal. For example, this can be achieved by indicating the terminal identification information (e.g., terminal serial number) of the corresponding target interactive terminal. It should also be noted that in some application scenarios, if there is only one preset interactive terminal, or if all preset interactive terminals need to be woken up simultaneously, the aforementioned terminal wake-up flag information can use a value of 0 or 1 to indicate whether to wake up the corresponding interactive terminal.

[0037] In this embodiment of the application, determining the terminal wake-up flag information based on the aforementioned interactive voice includes: The above-mentioned interactive voice is processed into text to obtain the interactive text corresponding to the interactive voice. Based on the interactive text, the terminal wake-up flag information is determined by a preset text processing model.

[0038] The text processing model described above can be set up in the cloud or locally on the interactive management terminal. In this embodiment, the text processing model is set up in the cloud to reduce the data storage and computational load on the interactive management terminal, and to utilize the more powerful computing capabilities of the cloud to improve processing efficiency.

[0039] Specifically, the aforementioned text processing model is deployed on a cloud processing platform; The above-mentioned interactive voice is processed into text to obtain the interactive text corresponding to the interactive voice. Based on the interactive text, the terminal wake-up flag information is determined through a preset text processing model, including: The aforementioned interactive voice is uploaded to the cloud processing platform to trigger the cloud processing platform to perform text conversion processing on the aforementioned interactive voice, obtain the interactive text corresponding to the aforementioned interactive voice, determine the processing result based on the aforementioned interactive text through the aforementioned text processing model, and return the aforementioned processing result to the aforementioned interactive management terminal. The above processing result includes terminal activation flag information indicating the target interactive terminal; or, the above processing result includes emotion indication information and response text matching the content of the above interactive text, as well as terminal activation flag information that does not indicate any target interactive terminal.

[0040] It should be noted that the above emotion indication information is used to indicate at least one target expression in the preset expression library.

[0041] Furthermore, the above method also includes: if the terminal wake-up flag information does not indicate any target interactive terminal, then displaying the target expression and outputting audio according to the response text.

[0042] Considering the limited local computing resources of the terminal, the text processing model used for speech analysis in this embodiment is deployed on a cloud processing platform. This platform has powerful computing capabilities, enabling rapid processing of speech data. The interactive management terminal uploads locally stored recording files to the cloud processing platform via Wi-Fi or mobile network and initiates a processing request.

[0043] After receiving the interactive voice, the cloud processing platform first processes it through a speech recognition engine, converting the speech signal into corresponding interactive text. Then, the interactive text is input into a pre-defined text processing model, which can be optimized based on a large language model (such as the GPT series models) and possesses the ability to understand the semantics of the text content, recognize intent, and analyze sentiment.

[0044] The text processing model performs multi-dimensional analysis of interactive text: on the one hand, it identifies the user's core needs and intentions and determines whether the need requires a specific external terminal to respond (e.g., if the user says "turn on the TV to play a movie", it determines that the smart TV needs to be activated); on the other hand, it analyzes the emotional information contained in the text (e.g., if the user says "I'm in a bad mood today and want to listen to a song", it identifies the emotion of sadness).

[0045] The cloud-based processing platform generates processing results based on the analysis results of the text processing model and returns them to the interactive management terminal. The processing results fall into two categories: Scenario 1: If the user requires a response from an external terminal, the processing result includes terminal wake-up flag information indicating the external terminal. Scenario 2: If the user's request does not require an external terminal response (such as pure emotional exchange, information consultation, etc.), the processing result includes a terminal wake-up flag that does not indicate any target interactive terminal, as well as emotion indication information and response text that match the interactive text content. The emotion indication information is used to indicate the target emotion in the preset emotion library of the interactive management terminal (e.g., "sadness" corresponds to a "comforting expression"), and the response text is the answer content generated in response to the user's request.

[0046] In this embodiment of the application, the target interactive terminal is a smart TV; Sending the aforementioned interactive voice to the aforementioned target interactive terminal to trigger the target interactive terminal to interact with the aforementioned target object based on the aforementioned interactive voice includes: Send a voice assistant activation signal to the aforementioned smart TV to trigger the smart TV to activate the TV voice assistant; The aforementioned interactive voice is sent to the aforementioned TV voice assistant to trigger the aforementioned TV voice assistant to interact with the aforementioned target object based on the aforementioned interactive voice.

[0047] Specifically, the above-mentioned response to detecting an interactive wake word and acquiring the interactive voice input from the target object includes: In response to the detection of an interactive wake word, the recording function of the aforementioned interactive management terminal is activated to record the interactive voice input by the target object, generate an audio file, and store it locally. Sending the aforementioned interactive voice to the aforementioned TV voice assistant to trigger the TV voice assistant to interact with the aforementioned target object based on the aforementioned interactive voice includes: The aforementioned audio file is sent to the aforementioned TV voice assistant to trigger the TV voice assistant to perform voice interaction with the aforementioned target object based on the content of the aforementioned audio file.

[0048] In this embodiment, the interactive management terminal stores the interactive voice locally, reducing the risk of data loss. A pre-established communication connection (e.g., Bluetooth connection) is established between the interactive management terminal and the smart TV to ensure stable data transmission. During data transmission, the recording file is directly sent to the TV voice assistant. The TV voice assistant receives the recording file, parses and processes it, extracts the user's voice commands, and executes corresponding operations based on the commands (such as playing a movie or adjusting the volume). Simultaneously, it provides feedback to the user through the TV screen or speakers, completing the interaction with the user.

[0049] If the terminal's activation flag does not indicate any target interactive terminal, the interactive management terminal will automatically complete the interaction with the user. Specifically, the terminal retrieves the corresponding target emoticon from a preset emoticon library based on the emotion indication information in the processing result and displays it on its own screen. Simultaneously, it inputs the response text into a speech synthesis engine, converts it into a speech signal, and outputs it through a speaker, achieving emotional interaction with the user. For example, if the user says, "I'm unhappy today," the interactive management terminal displays a frowning, comforting emoticon and reads the response text, "Don't be sad, let me chat with you."

[0050] In this way, when engaging in voice interaction, users do not need to issue voice commands individually for each terminal device. Instead, they only need to wake up the interaction management terminal and input the interactive voice. The interaction management terminal will automatically determine the terminal wake-up flag information, thereby identifying the target interactive terminal to be activated, and sending the interactive voice to the target interactive terminal to trigger the interaction between the target terminal and the target object. Users only need to issue interactive voice to the interaction management terminal to link devices together and realize the invocation of other terminal devices for voice interaction. This simplifies the user's voice interaction operation process, improves the convenience of voice interaction, and enhances the user experience.

[0051] In this embodiment of the application, the above-mentioned voice interaction processing method is further described in detail based on some specific application scenarios. Figure 2 This is a schematic diagram illustrating a specific flow of a voice interaction processing method provided in an embodiment of this application, such as... Figure 2 As shown, after the interactive management device is powered on, it connects to the wireless network and enters a wake-up waiting state. Upon detecting a wake-up word, it starts recording, saving the audio data as a local recording file for later retrieval, and simultaneously sending it to the cloud. The cloud converts the audio into text and interprets it based on a large model to determine if TV processing is required. Based on this, it generates a specific processing result and returns it to the interactive management terminal. The interactive management terminal then performs further processing based on the received result, determining whether the TV voice assistant needs to be activated. If so, the recording file is sent to the TV voice assistant to trigger voice interaction; otherwise, relevant facial expressions are displayed on the screen, and the corresponding text information is read aloud through the speaker.

[0052] Figure 3 This is a schematic diagram of a voice data processing flow provided in an embodiment of this application, such as... Figure 3 As shown in this embodiment, audio is acquired through the microphone array of the interactive management terminal. The acquired data is encoded using an analog-to-digital converter (ADC) and transmitted via the integrated circuit's built-in audio bus (I2S). Combined with the AFE processing algorithm, AEC algorithm, and WakeNet model, interference from environmental noise and its own playback sound is effectively removed, and wake-up word triggering is achieved. Simultaneously, a call interface is provided to output to the upper layer, allowing for both local storage of audio data and transmission to the cloud for processing. After cloud processing, a corresponding flag is returned, determining whether it needs to be sent to remote devices such as smart TVs. If transmission is required, the local recording file is read, ADPCM encoded, and transmitted to the remote device via Bluetooth Low Energy (BLE). If the audio data does not need to be sent to other devices, the response text returned from the cloud is converted using a digital-to-analog converter (DAC), and the converted sound data is output through a speaker.

[0053] In this way, the interactive management terminal can make corresponding facial expressions to engage in emotional interaction based on the dialogue content, and connect with the TV voice assistant to control the TV based on the dialogue content. Based on a large model, the dialogue content can be analyzed to determine whether the content needs to be sent to the TV for processing; if so, the TV voice assistant will be activated to send the message.

[0054] It should be noted that this application's solution is compatible with all display devices (such as televisions) that support Bluetooth remote voice control. It connects to the television via Bluetooth for voice interaction. Voice recordings are stored locally. When information is returned from the cloud for processing by the television's voice assistant, the local recording file is read, and the television's voice assistant is activated via Bluetooth, sending the voice data to the television for processing. The large-scale model determines whether the content is to be processed by the television based on the text, and matches the text content with facial expressions from an emoji library to generate a response text, achieving emotional interaction. In one specific application scenario, the interactive management terminal connects to the network via Wi-Fi and enters a wake-up state, establishing a communication connection with the smart television via Bluetooth. When a user speaks the wake-up word to activate the interactive management terminal and asks about today's weather, the interactive management terminal saves the voice information to the memory card and sends it to the cloud to be converted into text and sent to the large model for processing. At this time, the large model interprets the text and determines that the voice content needs to be sent to the TV for processing. The large model sends the voice flag that needs to be sent to the TV for processing. The interactive management terminal recognizes the voice flag that needs to be sent to the TV for processing, activates the TV's voice assistant, encodes the local voice file and sends it to the TV. The TV processes the voice information, obtains the local weather and displays the report.

[0055] In another specific application scenario, a user inputs, "I'm unhappy today, can you tell me a story?" The interactive management terminal saves the voice information to a memory card and sends it to the cloud to be converted into text and sent to a large model for processing. The large model then interprets the text, determines that the voice does not need to be sent to the TV, identifies the speaker's emotion, sends a comforting emotion flag, and sends the corresponding text data to the interactive management terminal. The interactive management terminal recognizes the voice flag that does not need to be sent to the TV and the comforting emotion flag, and then plays the received text through the speaker while displaying a comforting expression on the screen.

[0056] In this way, by using the interactive management terminal as the core hub, the target interactive terminal is determined based on the user's voice interaction and the voice data transmission is completed. This breaks down the interaction barriers between different devices, enabling cross-device collaborative interaction centered on the interactive management terminal. This simplifies the user's operation process and improves interaction efficiency. For example, users can send control commands to a smart TV through the same interactive management terminal without having to operate the TV separately, making operation more convenient.

[0057] By introducing a cloud-based text processing model for in-depth analysis of interactive voice, this technology can not only accurately determine whether an external terminal needs to be invoked, but also recognize emotional information in the user's voice and match corresponding facial expressions with response text. This achieves emotional interactive feedback, effectively improving the humanization of the user experience and solving the problem of the lack of emotionality in existing voice assistant technologies. It also addresses the technical issue of TV voice assistants being unable to perform emotional interaction by establishing communication between the interactive management terminal and the voice interaction terminal (e.g., a TV) via Bluetooth data transmission, enabling interaction between devices.

[0058] After acquiring the interactive voice, it is generated as an audio file and stored locally. During subsequent transmission, the locally stored audio file is directly called, avoiding the problem of voice data loss due to network fluctuations, ensuring the stability of voice interaction and improving the reliability of the interaction process.

[0059] Furthermore, in this embodiment, the communication method between the interactive management terminal and the target interactive terminal (such as a smart TV) is simple and efficient, requiring no large-scale hardware modification of the target interactive terminal. It only needs to support the voice assistant activation and voice data reception functions, which is highly compatible and easy to promote and apply in the existing smart terminal ecosystem.

[0060] like Figure 4 As shown in the figure, corresponding to the above-described voice interaction processing method, this application embodiment also provides a voice interaction processing system, the voice interaction processing system comprising: The data acquisition module 410 is used to acquire the interactive voice input by the target object in response to the detection of an interactive wake word; Data processing module 420 is used to determine terminal wake-up flag information based on the above interactive voice, wherein the terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up. The interaction processing module 430 is used to send the interactive voice to the target interactive terminal if the terminal wake-up flag information indicates at least one target interactive terminal, so as to trigger the target interactive terminal to interact with the target object according to the interactive voice.

[0061] In this way, when engaging in voice interaction, users do not need to issue voice commands individually for each terminal device. Instead, they only need to wake up the interaction management terminal and input the interactive voice. The interaction management terminal will automatically determine the terminal wake-up flag information, thereby identifying the target interactive terminal to be activated, and sending the interactive voice to the target interactive terminal to trigger the interaction between the target terminal and the target object. Users only need to issue interactive voice to the interaction management terminal to link devices together and realize the invocation of other terminal devices for voice interaction. This simplifies the user's voice interaction operation process, improves the convenience of voice interaction, and enhances the user experience.

[0062] It should be noted that the specific structure and implementation of the above-mentioned voice interaction processing system and its various modules or units can be referred to the corresponding descriptions in the above method embodiments, and will not be repeated here.

[0063] It should be further noted that the division of the various modules in the above-mentioned voice interaction processing system is not unique and is not intended as a specific limitation.

[0064] Based on the above embodiments, this application also provides an interactive management terminal, the principle block diagram of which can be as follows: Figure 5 As shown. The aforementioned interactive management terminal includes a processor, memory, network interface, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps of any of the aforementioned voice interaction processing methods. The display screen of the terminal can be a liquid crystal display (LCD) or an e-ink display.

[0065] Those skilled in the art will understand that Figure 5 The block diagram shown is only a partial structural diagram related to the solution of this application and does not constitute a limitation on the terminal on which the solution of this application is applied. The specific terminal may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0066] Figure 6 This is a six-view diagram of an interactive management terminal provided in an embodiment of this application. It should be noted that... Figure 6 This is merely an illustration of the appearance of the interactive management terminal and is not intended as a specific limitation. Figure 7 This is a schematic diagram illustrating an application scenario of an interactive management terminal provided in an embodiment of this application, such as... Figure 7As shown, the aforementioned interactive management terminal can be set up on a smart TV to serve as a companion for the user, but the specific setup method and location of the interactive management terminal are not specifically limited. For example, in some application scenarios, the aforementioned interactive management terminal can also be a portable terminal.

[0067] In one embodiment, a terminal is provided, the terminal including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of any of the voice interaction processing methods provided in the embodiments of this application.

[0068] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the voice interaction processing methods provided in this application.

[0069] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0070] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0071] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0072] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0073] In the embodiments provided in this application, it should be understood that the disclosed systems / terminal devices and methods can be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units described above is merely a logical functional division, and in actual implementation, it can be divided in other ways. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0074] If the integrated modules / units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, and software distribution media, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.

[0075] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions are not in essence a departure from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A voice interaction processing method, characterized in that, The method is applied to an interactive management terminal, and the method includes: In response to the detection of an interactive wake word, the interactive voice input from the target object is obtained; Based on the interactive voice, terminal wake-up flag information is determined, wherein the terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up; If the terminal wake-up flag information indicates at least one target interactive terminal, the interactive voice is sent to the target interactive terminal to trigger the target interactive terminal to interact with the target object according to the interactive voice.

2. The voice interaction processing method according to claim 1, characterized in that, Determining the terminal wake-up flag information based on the interactive voice includes: The interactive voice is processed into text to obtain the interactive text corresponding to the interactive voice, and the terminal wake-up flag information is determined based on the interactive text through a preset text processing model.

3. The voice interaction processing method according to claim 2, characterized in that, The text processing model is deployed on a cloud processing platform; The step of performing text conversion processing on the interactive voice to obtain the interactive text corresponding to the interactive voice, and determining the terminal wake-up flag information based on the interactive text using a preset text processing model, includes: The interactive voice is uploaded to the cloud processing platform to trigger the cloud processing platform to perform text conversion processing on the interactive voice to obtain the interactive text corresponding to the interactive voice. Based on the interactive text, the processing result is determined by the text processing model and the processing result is returned to the interactive management terminal. The processing result includes terminal activation flag information indicating the target interactive terminal; or, the processing result includes emotion indication information and response text matching the content of the interactive text, as well as terminal activation flag information not indicating any target interactive terminal.

4. The voice interaction processing method according to claim 3, characterized in that, The emotion indication information is used to indicate at least one target expression in a preset expression library.

5. The voice interaction processing method according to claim 4, characterized in that, The method further includes: If the terminal wake-up flag information does not indicate any target interactive terminal, the target emoticon is displayed, and audio is output according to the response text.

6. The voice interaction processing method according to any one of claims 1 to 5, characterized in that, The target interactive terminal is a smart TV; Sending the interactive voice to the target interactive terminal to trigger the target interactive terminal to interact with the target object based on the interactive voice includes: Send a voice assistant activation signal to the smart TV to trigger the smart TV to activate the TV voice assistant; The interactive voice is sent to the TV voice assistant to trigger the TV voice assistant to interact with the target object based on the interactive voice.

7. The voice interaction processing method according to claim 6, characterized in that, The step of acquiring interactive voice input from the target object in response to detecting an interactive wake word includes: In response to the detection of an interactive wake word, the recording function of the interactive management terminal is activated to record the interactive voice input by the target object, generate an audio file, and store it locally. Sending the interactive voice to the TV voice assistant to trigger the TV voice assistant to interact with the target object based on the interactive voice includes: The audio file is sent to the TV voice assistant to trigger the TV voice assistant to perform voice interaction with the target object based on the content of the audio file.

8. A voice interaction processing system, characterized in that, The system is applied to an interactive management terminal, and the system includes: The data acquisition module is used to acquire the interactive voice input by the target object in response to the detection of an interactive wake word; The data processing module is used to determine terminal wake-up flag information based on the interactive voice, wherein the terminal wake-up flag information is used to indicate the target interactive terminal that needs to be woken up. An interaction processing module is configured to send the interactive voice to the target interactive terminal if the terminal wake-up flag information indicates at least one target interactive terminal, so as to trigger the target interactive terminal to interact with the target object according to the interactive voice.

9. An interactive management terminal, characterized in that, The interactive management terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the voice interaction processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the voice interaction processing method as described in any one of claims 1 to 7.