Audio data playing method and device based on voice interaction, storage medium and computer program product

By generating a playback address and establishing a TCP connection in advance when a voice interaction request is made, combined with streaming audio synthesis and compression encoding technology, the problem of long connection establishment time in the traditional voice interaction process is solved, fast and stable audio data transmission and playback are achieved, and the user experience is improved.

CN120729660APending Publication Date: 2025-09-30HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510872625.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the traditional voice interaction process, sending the URL + DNS resolution + requesting to establish a connection takes a long time, resulting in a long voice interaction time and affecting the user experience.

Method used

After receiving the voice interaction request from the terminal device, the TTS service of the cloud server is immediately called to generate a playback address, and it is sent to the terminal device to establish a TCP connection, transmit audio data in real time, and optimize data transmission through streaming audio synthesis, compression encoding and TRUNK protocol. The terminal device plays the audio data immediately after receiving the playback instruction.

Benefits of technology

It reduces the duration of voice interaction and improves user experience, especially in complex network environments and with personalized user needs, ensuring fast and stable transmission and playback of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729660A_ABST
    Figure CN120729660A_ABST
Patent Text Reader

Abstract

The invention discloses an audio data playing method and device based on voice interaction, a storage medium and a computer program product, and relates to the field of smart home. A TTS service in a cloud server is started to be called, a playing address is generated according to expected response content corresponding to the voice interaction request, and the playing address is used for indicating an address for storing audio data corresponding to the expected response content; sending the playing address to the terminal equipment to indicate the terminal equipment to establish TCP connection with the cloud server according to the playing address; and when it is determined that the audio data is stored in the playing address, sending a playing instruction to the terminal device to instruct the terminal device to play the audio data through the TCP connection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart home, and more specifically, to a method and device for playing audio data based on voice interaction, a storage medium, and a computer program product. Background Art

[0002] With the rapid development of artificial intelligence, human-computer voice interaction is becoming increasingly widely adopted as the primary entry point and control interaction method in the smart home field. This, in turn, places higher demands on intelligent voice interaction. As voice interaction is the primary method of product interaction, the duration of the interaction becomes crucial, as it directly impacts the user experience. If the duration of voice interaction is too long, users may become impatient, affecting the user experience.

[0003] The traditional full-link process of voice interaction starts with user awakening: the user says the wake-up word to wake up the smart home appliance voice interaction system; voice recognition: the smart home appliance system receives the user's voice input and sends it to the cloud from the terminal for voice recognition and converts the voice into text; semantic understanding: the cloud system performs semantic understanding of the text input by the user, analyzes the user's intentions and needs, and the system retrieves information based on the user's intentions and needs, selects the appropriate answer method, or performs corresponding actions to feedback to the terminal; voice synthesis: the terminal obtains the broadcast text through analysis and assembly and then sends it to the cloud for TTS synthesis.

[0004] In the speech interaction process described above, during the speech synthesis phase, RTOS devices currently synthesize speech in the cloud. After the resource URL is sent to the terminal, the terminal's TTS player uses DNS to locate the actual server IP address and establishes a data connection with the server. The audio is then transmitted to the player via this data connection, and the player finally begins playing. The process of sending the URL, resolving the DNS, and establishing the connection takes an average of over 200ms, resulting in long wait times for users and a poor interaction experience.

[0005] Regarding the related technologies, in the traditional voice interaction process, sending the URL + DNS resolution + requesting to establish a connection takes a long time, resulting in the problem of excessively long voice interaction time. No effective solution has yet been proposed. Summary of the Invention

[0006] The embodiments of the present application provide a method and device for playing audio data based on voice interaction, a storage medium, and a computer program product, so as to at least solve the problem in the prior art that in the traditional voice interaction process, sending a URL+DNS resolution+requesting to establish a connection takes a long time, resulting in excessively long voice interaction time.

[0007] According to one embodiment of the present application, a method for playing audio data based on voice interaction is provided, including: when a terminal device initiates a voice interaction request, starting to call the TTS service in the cloud server to generate a playback address according to the expected response content corresponding to the voice interaction request, wherein the playback address is used to indicate the address of storing the audio data corresponding to the expected response content; sending the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address; when it is determined that the audio data is stored at the playback address, sending a playback instruction to the terminal device to instruct the terminal device to play the audio data through the TCP connection.

[0008] In an exemplary embodiment, when a voice interaction request is monitored from a terminal device, the TTS service in the cloud server is called to generate a playback address according to the expected response content corresponding to the voice interaction request, including: when a voice signal corresponding to the voice interaction request is monitored, the TTS service is called; the content type and content format of the expected response content are determined according to the voice interaction request, wherein the expected response content corresponds to the audio data; and the playback address is generated through the TTS service according to the content type and the content format.

[0009] In an exemplary embodiment, when it is determined that the audio data is stored in the playback address, a playback instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection, including: generating an audio data stream of the audio data through streaming audio synthesis technology, and storing the audio data stream to the playback address in real time; when it is detected that the audio data stream starts to be stored in the playback address, compressing and encoding the audio data stream to obtain a compressed data packet, and sending the playback instruction to the terminal device to instruct the terminal device to obtain the compressed data packet through the TCP connection, and play the audio data corresponding to the compressed data packet according to the playback instruction, wherein the playback instruction carries metadata of the audio data stream, and the metadata includes: audio format, sampling rate and channel configuration.

[0010] In an exemplary embodiment, the method further includes: obtaining the decoding capability of the terminal device and the real-time network bandwidth of the TCP connection; and adjusting the data format of the compressed data packet in real time according to the decoding capability and the real-time network bandwidth.

[0011] In an exemplary embodiment, after the audio data stream is compressed and encoded to obtain a compressed data packet, the method further includes: transmitting the compressed data packet to the terminal device in real time based on the TCP connection through the TRUNK protocol; sending the play instruction to the terminal device to instruct the terminal device to cache the audio data corresponding to the compressed data packet when receiving the compressed data packet, and play the cached audio data in real time.

[0012] In an exemplary embodiment, the method further includes: analyzing historical requests of the target object through a machine learning algorithm, and generating a first user portrait of the target object based on the analysis results, wherein the first user portrait is at least used to indicate the living habits of the target object; and sending the target playback address corresponding to the living habits to the terminal device at the target time point according to the first user portrait, so as to instruct the terminal device to pre-establish the TCP connection at the target time point.

[0013] In an exemplary embodiment, before sending the play instruction to the terminal device, the method further includes: obtaining object information of the target object from the cloud server, and generating a second user portrait of the target object based on the object information, wherein the second user portrait is used to indicate the voice interaction habits of the target object; determining sound attribute information for voice interaction based on the voice interaction habits, wherein the sound attribute information includes: timbre, speaking speed, and tone; adding the sound attribute information to the play instruction to instruct the terminal device to play the audio data according to the sound attribute information.

[0014] According to another embodiment of the embodiments of the present application, a device for playing audio data based on voice interaction is also provided, including: a generation module, which is used to start calling the TTS service in the cloud server to generate a playback address according to the expected response content corresponding to the voice interaction request when monitoring a voice interaction request initiated by a terminal device, wherein the playback address is used to indicate the address of the audio data corresponding to the expected response content; an establishment module, which is used to send the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address; and a sending module, which is used to send a playback instruction to the terminal device when determining that the audio data is stored at the playback address to instruct the terminal device to play the audio data through the TCP connection.

[0015] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for playing audio data based on voice interaction when running.

[0016] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in each embodiment of the present application when executed by a processor.

[0017] Through this application, after receiving a voice interaction request initiated by a terminal device, the TTS service in the cloud server is immediately called to generate a playback address according to the expected response content corresponding to the voice interaction request, wherein the playback address is used to store the audio data corresponding to the voice interaction request; the playback address is then sent to the terminal device to instruct the terminal device to establish a TCP connection according to the playback address; after determining that the audio data is stored in the playback address, a play instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection; using the above solution, a URL is sent in advance when starting voice interaction, and the terminal establishes a data connection with the cloud in advance. When the audio data stream needs to be pushed later, the link is directly used to start streaming playback, thereby reducing the voice interaction time; thereby solving the problem in the related technology that in the traditional voice interaction process, sending a URL + DNS resolution + requesting to establish a connection takes a long time, resulting in excessively long voice interaction time. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 This is a hardware structure block diagram of a cloud server for a method for playing audio data based on voice interaction according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of a method for playing audio data based on voice interaction according to an embodiment of the present application;

[0022] Figure 3 This is a structural block diagram of a device for playing audio data based on voice interaction according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] The method embodiments provided in the embodiments of the present application can be executed in a cloud server or similar computing device. Taking running on a cloud server as an example, Figure 1 This is a hardware structure block diagram of a cloud server for a method of playing audio data based on voice interaction according to an embodiment of the present application. Figure 1 As shown, the cloud server may include one or more ( Figure 1 Only one is shown) processor 102 (processor 102 may include but is not limited to a microprocessor (Central Processing Unit, MCU) or a programmable logic device (Field Programmable Gate Array, FPGA) and a memory 104 for storing data, wherein the cloud server may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the cloud server. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0026] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for playing audio data based on voice interaction in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to a cloud server via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0027] The cloud server's communications provider provides a wireless network. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one embodiment, the transmission device 106 can be a radio frequency (RF) module for wireless communication with the Internet.

[0028] In this embodiment, a method for playing audio data based on voice interaction is provided. Figure 2 is a flowchart of a method for playing audio data based on voice interaction according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps S202-S206:

[0029] Step S202: When a voice interaction request is detected by the terminal device, the TTS service in the cloud server is called to generate a playback address based on the expected response content corresponding to the voice interaction request, where the playback address indicates the address where the audio data corresponding to the expected response content is stored;

[0030] Step S204: sending the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address;

[0031] Step S206: When it is determined that the audio data is stored at the playback address, a playback instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection.

[0032] Through the above steps, after receiving the voice interaction request initiated by the terminal device, the TTS service in the cloud server is immediately called to generate a playback address according to the expected response content corresponding to the voice interaction request, wherein the playback address is used to store the audio data corresponding to the voice interaction request; the playback address is then sent to the terminal device to instruct the terminal device to establish a TCP connection according to the playback address; after determining that the audio data is stored in the playback address, a play instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection; using the above solution, a URL is sent in advance when starting voice interaction, and the terminal establishes a data connection with the cloud in advance. When the audio data stream needs to be pushed later, the link is directly used to start streaming playback, thereby reducing the voice interaction time; thereby solving the problem in related technologies that in the traditional voice interaction process, sending URL+DNS resolution+requesting to establish a connection takes a long time, resulting in excessively long voice interaction time.

[0033] Optionally, when a voice interaction request is monitored from a terminal device, the TTS service in the cloud server is called to generate a playback address according to the expected response content corresponding to the voice interaction request, including: when a voice signal corresponding to the voice interaction request is monitored, the TTS service is called; the content type and content format of the expected response content are determined according to the voice interaction request, wherein the expected response content corresponds to the audio data; and the playback address is generated through the TTS service according to the content type and the content format.

[0034] This example mainly describes how to generate a playback address in advance by calling the TTS service of a cloud server when a voice interaction request occurs. The core of this process is to accelerate the preparation of TTS synthesis and audio playback by predicting the expected response content of the user request, thereby improving the overall interaction experience. The following is a detailed description of the steps of this example:

[0035] A. Listening for Voice Interaction Requests: When a terminal device (such as a smart speaker or mobile app) is in standby mode, the system continuously listens for voice signals from the user. Once a valid voice trigger (such as a wake-up word) is detected, the system considers it to have received a voice interaction request initiated by the user.

[0036] B. Invoke the TTS service: As soon as the system detects a voice signal, it immediately initiates a call to the TTS service in the cloud server. This means the system no longer passively waits for speech recognition and semantic understanding to complete, but instead proactively prepares for the expected response content.

[0037] C. Determine the expected response content: When invoking the TTS service, the system predicts and determines the expected response content based on the user's voice interaction request. This step includes determining the content type (such as weather query, music playback, smart home control feedback, etc.) and content format (such as MP3, AAC, OPUS, and other audio formats).

[0038] Content type prediction: By analyzing the user's speech patterns, historical request records, and scene context information, the system can predict the type of information the user may request. For example, if the user says "What's the weather like today", the system predicts that this is a weather query request.

[0039] Content format determination: Select the most appropriate content format based on the audio playback capabilities of the terminal device and the network environment to ensure that the audio stream can be transmitted efficiently and losslessly.

[0040] D. Generate a playback address: Based on the predicted content type and determined content format, the system generates a corresponding playback address through the TTS service. The generated playback address is a network identifier pointing to the cloud audio resource. It contains information about the expected response content and the specified audio format, ensuring that the terminal device can accurately and quickly address and start playback based on this address. The system then integrates the expected audio resource, format parameters, and possible encryption information into a playback address. This address is the key for the terminal to connect to the cloud and obtain the audio stream.

[0041] E. Pre-preparation: Before speech recognition and semantic understanding are fully completed, the terminal device has already received the playback address. The terminal can start DNS resolution, pre-establish a TCP connection, and prepare the player. When the TTS-synthesized audio data stream arrives, it can immediately start playback without any additional waiting time.

[0042] F. Real-time playback: Once the cloud completes TTS synthesis, the system can directly notify the terminal device that the audio data is ready. The terminal device immediately starts streaming and playing according to the pre-acquired playback address, achieving truly seamless connection and real-time response.

[0043] This implementation significantly reduces the latency from user request to response by front-loading the TTS service call and combining it with prediction and optimization of expected response content, thus improving the user experience. It also demonstrates how this technical solution can flexibly adapt to complex network environments and diverse user needs in practical applications.

[0044] Optionally, when it is determined that the audio data is stored in the playback address, a playback instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection, including: generating an audio data stream of the audio data through streaming audio synthesis technology, and storing the audio data stream to the playback address in real time; when it is detected that the audio data stream starts to be stored in the playback address, compressing and encoding the audio data stream to obtain a compressed data packet, and sending the playback instruction to the terminal device to instruct the terminal device to obtain the compressed data packet through the TCP connection, and play the audio data corresponding to the compressed data packet according to the playback instruction, wherein the playback instruction carries metadata of the audio data stream, and the metadata includes: audio format, sampling rate and channel configuration.

[0045] This example describes how, in a voice interaction TTS pre-establishment strategy, a cloud server can efficiently and intelligently send playback instructions to a terminal device once the audio data is ready. The following steps enable fast playback and optimized transmission of audio data:

[0046] After the speech recognition and semantic understanding phases are complete, the audio data needed to respond to the user's request (e.g., a spoken message) is determined. At this point, the cloud server invokes the TTS (Text-to-Speech) service to begin generating the audio data stream. This process utilizes streaming audio synthesis technology, allowing audio data to be generated incrementally during the synthesis process, without having to wait until the entire audio segment is fully synthesized.

[0047] The audio data stream is generated and stored in real time to a pre-generated playback address. This address is usually a temporary resource URL used to carry the audio stream to be played.

[0048] The cloud server continuously monitors the storage status of the audio data stream at the playback address. Once the system detects the start of audio data stream storage, it immediately performs compression encoding. Compression encoding is designed to accelerate audio data transmission and reduce data volume, especially in environments with unstable network conditions or limited bandwidth. Compressed audio data packets can be transmitted more quickly to the terminal device via the TCP connection.

[0049] During the compression encoding process, the cloud considers factors such as audio format compatibility, sampling rate rationality, and channel configuration matching to generate compressed data packets. This step not only improves transmission efficiency but also ensures that the terminal device can correctly decode and play the audio data.

[0050] The cloud server sends a play command to the terminal device. This command contains metadata about the compressed data packet, including key information such as the audio format, sampling rate, and channel configuration. Upon receiving the play command, the terminal device immediately decodes and plays the audio data based on this metadata, without requiring additional configuration or interpretation, thus achieving rapid audio response and playback.

[0051] The terminal device receives the compressed data packet through the established TCP connection and starts playing the audio data. The TCP connection ensures the reliability of audio data transmission and can ensure the complete delivery of audio data even under poor network conditions.

[0052] In summary, this embodiment, through streaming audio synthesis, real-time storage, compression encoding processing, and the transmission of playback instructions with metadata, enables terminal devices to quickly respond to and play audio data synthesized by cloud servers, significantly improving the smoothness of voice interaction and user experience. Furthermore, by optimizing audio data transmission, it ensures high efficiency and stability in various network environments, demonstrating the innovative and practical nature of this technical solution.

[0053] Optionally, the method further includes: obtaining the decoding capability of the terminal device and the real-time network bandwidth of the TCP connection; and adjusting the data format of the compressed data packet in real time according to the decoding capability and the real-time network bandwidth.

[0054] This embodiment further optimizes the audio data transmission process in the "Voice Interaction TTS Pre-establishment Strategy and Method" by dynamically adjusting the audio data compression format based on the decoding capabilities of the terminal device and the real-time network bandwidth of the TCP connection, thereby achieving more efficient and stable transmission and playback. The details are as follows:

[0055] At the start of a voice interaction, the cloud server first obtains information about the terminal's decoding capabilities. This may include the terminal's support for specific audio formats (such as MP3, AAC, Opus, etc.), the performance of the decoder (such as the maximum supported bit rate and sampling rate), and the terminal's memory and processor resources. This information helps the cloud determine the audio compression format and parameters that are most suitable for the terminal.

[0056] The cloud server also continuously monitors the real-time network bandwidth of the TCP connection between the cloud server and the end device. Fluctuations in network bandwidth can affect the transmission efficiency of audio data, especially in mobile networks or unstable Wi-Fi environments. Real-time bandwidth monitoring helps the cloud server make more informed decisions during transmission.

[0057] Based on the information obtained above, the cloud server will adjust the data format of the compressed data packet in real time during the process of generating and compressing the audio data stream. For example: if the decoder performance of the terminal device is weak, the cloud may choose a lower bit rate audio format for compression to reduce the decoding burden and ensure that the terminal device can play smoothly. If the real-time network bandwidth is low, the cloud will automatically select an audio format with a higher compression ratio and smaller data volume. Even in poor network conditions, it can ensure fast transmission of audio data and reduce playback delays. When the network bandwidth is sufficient and the terminal decoding capability is strong, the cloud can generate high-quality audio data packets to enhance the user's listening experience.

[0058] Through this embodiment, the cloud server can intelligently adjust the audio data transmission strategy based on the actual capabilities of the terminal device and network conditions, ensuring optimal audio playback in different scenarios. This dynamic adjustment mechanism not only improves the real-time nature of voice interaction, but also enhances the system's adaptability and user experience. This is especially true when processing high-density, high-quality audio data, effectively avoiding playback issues and transmission bottlenecks caused by decoding capabilities and network bandwidth limitations.

[0059] This technical solution reflects in-depth consideration of users' personalized needs and complex network environments, and is an important innovation for improving the performance of voice interaction systems and user satisfaction.

[0060] In an exemplary embodiment, after the audio data stream is compressed and encoded to obtain a compressed data packet, the method further includes: transmitting the compressed data packet to the terminal device in real time based on the TCP connection through the TRUNK protocol; sending the play instruction to the terminal device to instruct the terminal device to cache the audio data corresponding to the compressed data packet when receiving the compressed data packet, and play the cached audio data in real time.

[0061] This example describes in detail how, in the "Voice Interaction TTS Pre-establishment Strategy and Method," a cloud server compresses and encodes the audio data stream, transmits the compressed data packet to the terminal device in real time through a combination of the TRUNK protocol and TCP connection, and instructs the terminal device to cache and play the data in real time. The details are as follows:

[0062] The cloud server compresses and encodes the audio data stream generated by the TTS service into compressed data packets. The purpose of compression encoding is to reduce the amount of audio data transmitted and speed up transmission while ensuring that the audio quality is not significantly affected.

[0063] The TRUNK protocol allows for the transmission of multiple data types (such as audio, video, and control information) over a single connection. In this embodiment, the cloud server uses the TRUNK protocol to transmit compressed data packets to the terminal device in real time over an established TCP connection. The use of the TRUNK protocol enables the simultaneous transmission of audio data and communication of control commands over the same TCP connection, improving connection utilization and transmission efficiency.

[0064] The compressed data packets transmitted via the TRUNK protocol ensure that the audio data, after being synthesized in the cloud, reaches the terminal device quickly via the TCP connection. Because the TCP connection is established during the speech recognition phase, this step allows for delay-free data transmission.

[0065] After receiving the compressed data packets, the terminal device caches them. The purpose of caching is to smooth out playback interruptions caused by network fluctuations and ensure that the terminal player can play audio smoothly even when the network conditions are poor. In addition, after the terminal device receives the playback command sent by the cloud, it will immediately begin decompressing and playing the audio data in real time based on the metadata carried by the command (such as audio format, sampling rate, channel configuration, etc.) and the caching strategy.

[0066] This embodiment achieves efficient real-time transmission of audio data by combining the TRUNK protocol with TCP connection. At the same time, through the caching mechanism of the terminal device, it ensures the continuity and stability of audio playback, providing a good user experience even when network conditions change.

[0067] This technology is particularly suitable for scenarios with high real-time requirements, such as voice responses from smart assistants, voice courses in online education, and voice synchronization in video conferencing. It can significantly improve the response speed and fluency of voice interaction, reduce user waiting time, and enhance the reliability of real-time voice communications.

[0068] This embodiment demonstrates how to achieve high-quality, low-latency real-time voice communication in a voice interaction system by optimizing audio transmission protocols and processes. It is one of the key technical solutions to improve the performance of voice interaction systems.

[0069] In an exemplary embodiment, the method further includes: analyzing historical requests of the target object through a machine learning algorithm, and generating a first user portrait of the target object based on the analysis results, wherein the first user portrait is at least used to indicate the living habits of the target object; and sending the target playback address corresponding to the living habits to the terminal device at the target time point according to the first user portrait, so as to instruct the terminal device to pre-establish the TCP connection at the target time point.

[0070] This embodiment introduces machine learning algorithms and user profiling concepts to further optimize the "Voice Interaction TTS Pre-Connection Strategy and Method," enabling the cloud server to predictively push playback addresses to users and establish TCP connections in advance, significantly shortening the response time of voice interaction. The following is a detailed description:

[0071] First, the cloud server uses machine learning algorithms to conduct an in-depth analysis of the historical request data of the target object (i.e., the user). Historical request data may include details such as the frequency of users using voice interaction in different time periods, the types of content frequently inquired about (such as news, weather, music, etc.), the preferred voice broadcast timbre, speaking speed, etc. By analyzing this data, the system can construct the user's behavior patterns and living habits, thereby generating the user's first user profile. The user profile not only reflects the user's basic information and preferences, but more importantly, it can capture the user's living habits, such as checking the weather forecast at a fixed time every day, listening to the news during the weekday commute, etc.

[0072] Based on the generated first user profile, the cloud server can predict the user's likely voice requests at a specific time and generate a target playback URL accordingly. For example, if the user profile indicates that the user asks for the weather forecast every morning at 7:00 AM, the cloud server will generate a playback URL for the weather forecast before that time. The advantage of this is that the system has ample time to prepare and optimize audio data transmission before the actual voice request arrives.

[0073] Once the target playback address is determined, the cloud server will send it to the terminal device at the predicted target time. Unlike traditional methods that establish a connection only after a voice interaction request is made, this method pre-establishes a TCP connection immediately after the terminal device receives the playback address. This connection is used directly to transmit audio data when the user actually initiates a voice request, eliminating the time spent waiting for a connection to be established and significantly improving the speed and efficiency of voice interaction.

[0074] This technical solution not only solves the latency issues inherent in existing technologies due to steps like DNS resolution and connection establishment, but also introduces predictive and personalized elements, making voice interaction more tailored to users' actual needs and providing a smoother, more timely service experience. This technology is particularly suitable for scenarios like smart homes and personal assistants, proactively providing services tailored to users' lifestyles and habits, improving the intelligence and user satisfaction of voice assistants.

[0075] In summary, this embodiment uses machine learning to analyze user historical behavior, generate user portraits, predictively push playback addresses based on the portraits, and pre-establish a TCP connection, thereby greatly improving the response speed of the voice interaction system and significantly improving the user experience. It is an important direction for the development of modern smart home and smart voice assistant technologies.

[0076] In an exemplary embodiment, before sending the play instruction to the terminal device, the method further includes: obtaining object information of the target object from the cloud server, and generating a second user portrait of the target object based on the object information, wherein the second user portrait is used to indicate the voice interaction habits of the target object; determining sound attribute information for voice interaction based on the voice interaction habits, wherein the sound attribute information includes: timbre, speaking speed, and tone; adding the sound attribute information to the play instruction to instruct the terminal device to play the audio data according to the sound attribute information.

[0077] This embodiment deepens the personalized service features of the "Voice Interaction TTS Pre-establishment Strategy and Method" by obtaining the target object (user) information from the cloud server and then generating a second user profile that reflects the user's voice interaction preferences. This allows for customized adjustment of the voice properties of the speech synthesis before sending the playback command to the terminal device, providing an interactive experience that better suits the user's preferences. The following is a detailed description:

[0078] The cloud server first collects information about the target user. This information may come from multiple sources, including but not limited to the basic information provided by the user when registering an account, previous voice interaction records, and preferences set by the user through the app or other interfaces. This information covers the user's voice interaction habits, such as preferred voice type (male, female, child voice, etc.), preferred speaking speed (fast, medium, slow), and commonly used languages ​​and dialects.

[0079] Based on the collected subject information, the cloud uses data analysis and intelligent algorithms to generate a second user profile for the target subject. This profile not only reflects the user's static attributes (such as gender, age, occupation, etc.), but more importantly, captures the user's dynamic preferences and habits, especially personalized needs related to voice interaction.

[0080] Based on the voice interaction habits revealed by the second user profile, the cloud system determines a series of voice attribute information for voice interaction. This information includes key parameters such as voice color, speech rate, and tone of voice. These directly determine whether the TTS-synthesized speech sounds natural and friendly, and whether it can better convey information and emotions.

[0081] For example, if a user prefers a gentler female voice, a moderate speaking speed, and an encouraging and positive tone, the cloud server will adjust the voice attributes of the TTS service based on these preferences to make it more in line with the user's preferences.

[0082] Finally, before sending a play command to the terminal device, the cloud server embeds the determined sound attribute information into the play command. This means that when the terminal receives the play command, it not only knows which audio file to play, but also obtains detailed instructions on how to play it (i.e., the sound attributes).

[0083] The terminal device adjusts its internal audio playback parameters based on the sound attribute information such as timbre, speaking speed, and tone contained in the playback command, ensuring that the audio data is played in the user's preferred manner, providing a highly personalized voice interaction experience.

[0084] This technical solution not only solves the fundamental issue of voice interaction duration but also further incorporates personalized service concepts, enhancing user interactivity and satisfaction with the system. This technology is particularly suitable for scenarios requiring frequent voice interaction, such as smart home control, smart wearable devices, and virtual assistants. It can provide more attentive and natural voice services tailored to each user's unique preferences, making voice interaction more engaging and efficient.

[0085] In summary, this embodiment has taken the "voice interaction TTS pre-connection strategy and method" to a new level through precise user portrait analysis and personalized setting of voice attributes. It not only improves the practicality and innovation of the technology, but also brings users a more personalized voice interaction experience.

[0086] The problem with the existing technology is that after the speech synthesis is completed, the terminal playback URL is pushed. The terminal first performs DNS resolution on the IP, establishes a connection, and reads the audio for playback. Sending the URL + DNS resolution + requesting to establish a connection takes an average of 200ms+, which accounts for a large delay in the overall interaction process.

[0087] The audio data playback method provided by the above embodiment sends the URL to the terminal in advance during the voice recognition process. The terminal has enough time to perform DNS resolution, start the player and establish a TCP connection in advance. After the recognition is completed, the semantic understanding and the voice synthesis audio are inserted into the URL playback resource, the player can be directly notified to pull the stream for playback.

[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0089] In this embodiment, a playback device for audio data based on voice interaction is also provided. The playback device for audio data based on voice interaction is used to implement the above-mentioned embodiments and preferred embodiments, and the details that have been described will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0090] Figure 3 is a structural block diagram of an optional device for playing audio data based on voice interaction according to an embodiment of the present application; Figure 3 As shown, including:

[0091] A generating module 32 is configured to, upon monitoring a voice interaction request initiated by a terminal device, initiate invoking the TTS service in the cloud server to generate a playback address based on the expected response content corresponding to the voice interaction request, wherein the playback address indicates an address for storing audio data corresponding to the expected response content;

[0092] An establishing module 34 is configured to send the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address;

[0093] The sending module 36 is configured to, when determining that the audio data is stored at the playback address, send a playback instruction to the terminal device to instruct the terminal device to play the audio data through the TCP connection.

[0094] Through the above embodiment, after receiving the voice interaction request initiated by the terminal device, the TTS service in the cloud server is immediately called to generate a playback address according to the expected response content corresponding to the voice interaction request; the playback address is then sent to the terminal device to instruct the terminal device to establish a TCP connection according to the playback address, wherein TCP is used to transmit audio data; after determining that the audio data is stored in the playback address, a play instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection; using the above scheme, a URL is issued in advance when starting the voice interaction, and the terminal establishes a data connection with the cloud in advance. When the audio data stream needs to be pushed later, the link is directly used to start the streaming playback, thereby reducing the voice interaction time; thereby solving the problem in the related technology that in the traditional voice interaction process, issuing the URL+DNS resolution+requesting to establish a connection takes a long time, resulting in excessively long voice interaction time.

[0095] Optionally, the above-mentioned generation module 32 is also used to start calling the TTS service when a voice signal corresponding to the voice interaction request is monitored; determine the content type and content format of the expected response content according to the voice interaction request, wherein the expected response content corresponds to the audio data; and generate the playback address through the TTS service according to the content type and the content format.

[0096] Optionally, the above-mentioned sending module 36 is also used to generate an audio data stream of the audio data through streaming audio synthesis technology, and store the audio data stream to the playback address in real time; when it is detected that the audio data stream starts to be stored in the playback address, the audio data stream is compressed and encoded to obtain a compressed data packet, and the playback instruction is sent to the terminal device to instruct the terminal device to obtain the compressed data packet through the TCP connection, and play the audio data corresponding to the compressed data packet according to the playback instruction, wherein the playback instruction carries metadata of the audio data stream, and the metadata includes: audio format, sampling rate and channel configuration.

[0097] Optionally, the above-mentioned device is further used to obtain the decoding capability of the terminal device and the real-time network bandwidth of the TCP connection; and adjust the data format of the compressed data packet in real time according to the decoding capability and the real-time network bandwidth.

[0098] Optionally, the above-mentioned sending module 36 is also used to transmit the compressed data packet to the terminal device in real time based on the TCP connection through the TRUNK protocol; and send the play instruction to the terminal device to instruct the terminal device to cache the audio data corresponding to the compressed data packet when receiving the compressed data packet, and play the cached audio data in real time.

[0099] Optionally, the above-mentioned establishment module 34 is also used to analyze the historical requests of the target object through a machine learning algorithm, and generate a first user portrait of the target object based on the analysis results, wherein the first user portrait is at least used to indicate the living habits of the target object; according to the first user portrait, the target playback address corresponding to the living habits is sent to the terminal device at the target time point to instruct the terminal device to pre-establish the TCP connection at the target time point.

[0100] Optionally, the above-mentioned sending module 36 is also used to obtain object information of the target object from the cloud server, and generate a second user portrait of the target object based on the object information, wherein the second user portrait is used to indicate the voice interaction habits of the target object; determine the sound attribute information used for voice interaction based on the voice interaction habits, wherein the sound attribute information includes: timbre, speaking speed, and tone; add the sound attribute information to the playback instruction to instruct the terminal device to play the audio data according to the sound attribute information.

[0101] An embodiment of the present application further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.

[0102] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:

[0103] S1, when a voice interaction request is detected by a terminal device, starts calling the TTS service in the cloud server to generate a playback address according to the expected response content corresponding to the voice interaction request, wherein the playback address is used to indicate the address of the audio data corresponding to the expected response content;

[0104] S2, sending the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address;

[0105] S3: When it is determined that the audio data is stored at the playback address, a playback instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection.

[0106] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0107] An embodiment of the present application further provides a computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program product, and when the computer program is executed by a processor, the steps of the method described in each embodiment of the present application are implemented.

[0108] Optionally, in this embodiment, the computer program may be configured to implement the following steps when executed by a processor:

[0109] S1, when a voice interaction request is detected by a terminal device, starts calling the TTS service in the cloud server to generate a playback address according to the expected response content corresponding to the voice interaction request, wherein the playback address is used to indicate the address of the audio data corresponding to the expected response content;

[0110] S2, sending the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address;

[0111] S3: When it is determined that the audio data is stored at the playback address, a playback instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection.

[0112] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0113] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0114] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for playing audio data based on voice interaction, characterized in that: include: When a voice interaction request is detected by the terminal device, the TTS service in the cloud server is called to generate a playback address according to the expected response content corresponding to the voice interaction request, wherein the playback address is used to indicate the address of the audio data corresponding to the expected response content; Sending the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address; When it is determined that the audio data is stored at the playback address, a playback instruction is sent to the terminal device to instruct the terminal device to play the audio data through the TCP connection.

2. The method for playing audio data based on voice interaction according to claim 1, characterized in that: When a voice interaction request is detected by a terminal device, the TTS service in the cloud server is called to generate a playback address according to the expected response content corresponding to the voice interaction request, including: When a voice signal corresponding to the voice interaction request is monitored, starting to call the TTS service; determining a content type and a content format of the expected response content according to the voice interaction request, wherein the expected response content corresponds to the audio data; The playback address is generated by the TTS service according to the content type and the content format.

3. The method for playing audio data based on voice interaction according to claim 1, characterized in that: When it is determined that the audio data is stored at the playback address, sending a playback instruction to the terminal device to instruct the terminal device to play the audio data through the TCP connection, including: Generating an audio data stream of the audio data by using a streaming audio synthesis technology, and storing the audio data stream to the playback address in real time; When it is detected that the audio data stream starts to be stored in the playback address, the audio data stream is compressed and encoded to obtain a compressed data packet, and the playback instruction is sent to the terminal device to instruct the terminal device to obtain the compressed data packet through the TCP connection and play the audio data corresponding to the compressed data packet according to the playback instruction, wherein the playback instruction carries metadata of the audio data stream, and the metadata includes: audio format, sampling rate and channel configuration.

4. The method for playing audio data based on voice interaction according to claim 3, characterized in that: The method further comprises: Obtaining the decoding capability of the terminal device and the real-time network bandwidth of the TCP connection; The data format of the compressed data packet is adjusted in real time according to the decoding capability and the real-time network bandwidth.

5. The method for playing audio data based on voice interaction according to claim 3, characterized in that: After compressing and encoding the audio data stream to obtain a compressed data packet, the method further includes: transmitting the compressed data packet to the terminal device in real time based on the TCP connection via the TRUNK protocol; The play instruction is sent to the terminal device to instruct the terminal device to cache the audio data corresponding to the compressed data packet and play the cached audio data in real time when the compressed data packet is received.

6. The method for playing audio data based on voice interaction according to any one of claims 1 to 5, characterized in that: The method further comprises: Analyzing historical requests of a target object using a machine learning algorithm, and generating a first user profile of the target object based on the analysis results, wherein the first user profile is used to at least indicate the living habits of the target object; According to the first user portrait, the target playback address corresponding to the living habit is sent to the terminal device at the target time point to instruct the terminal device to pre-establish the TCP connection at the target time point.

7. The method for playing audio data based on voice interaction according to claim 3, characterized in that: Before sending the play instruction to the terminal device, the method further includes: Obtaining object information of a target object from the cloud server, and generating a second user profile of the target object based on the object information, wherein the second user profile is used to indicate the voice interaction habits of the target object; Determining voice attribute information for voice interaction based on the voice interaction habit, wherein the voice attribute information includes: timbre, speaking speed, and tone; The sound attribute information is added to the play instruction to instruct the terminal device to play the audio data according to the sound attribute information.

8. A device for playing audio data based on voice interaction, characterized in that: include: A generation module is configured to, upon monitoring a voice interaction request initiated by a terminal device, start calling the TTS service in the cloud server to generate a playback address based on the expected response content corresponding to the voice interaction request, wherein the playback address is used to indicate an address for storing audio data corresponding to the expected response content; An establishing module, configured to send the playback address to the terminal device to instruct the terminal device to establish a TCP connection with the cloud server according to the playback address; The sending module is used to send a play instruction to the terminal device when it is determined that the audio data is stored in the play address, so as to instruct the terminal device to play the audio data through the TCP connection.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the method according to any one of claims 1 to 7 is executed when the program is executed.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voice interaction implementation method and device, computer equipment and storage medium

    CN109637519A

  • Voice interaction method, voice interaction system, server and storage medium

    CN113421564A

  • Voice interaction method, server and computer storage medium

    CN115376511A

  • Test method, system and device of intelligent voice interaction system and medium

    CN118135998A

  • Voice interaction method, device, equipment, medium and program product

    CN120164464A