Information processing method and apparatus
Patent Information
- Application Number
- PCT/CN2026/070095
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-20
- Filing Date
- 2026-01-04
- Publication Date
- 2026-08-27
Smart Images

Figure CN2026070095_27082026_PF_FP_ABST
Abstract
Description
An information processing method and apparatus
[0001] This application claims priority to the following patent application, the entire contents of which are incorporated herein by reference.
[0002] 1. A Chinese patent application filed with the State Intellectual Property Office on February 20, 2025, with application number 2025101940915 and invention title "An Information Processing Method and Apparatus";
[0003] 2. A Chinese patent application filed with the State Intellectual Property Office on February 20, 2025, with application number 2025101923801 and invention title "A Communication Method and Device";
[0004] 3. A Chinese patent application filed with the State Intellectual Property Office on February 20, 2025, with application number 2025101932815 and invention title "A Communication Method and Device". Technical Field
[0005] This application relates to the field of communications, and in particular to an information processing method and apparatus. Background Technology
[0006] With the development of artificial intelligence (AI) technology, AI models, such as large AI models, are increasingly capable of understanding multimodal information. AI models can now understand not only text, but also images, audio, and even video. Correspondingly, the ways users interact in real-time with AI applications (apps) that deploy AI models are also changing. Currently, users and AI applications can engage in real-time voice conversations.
[0007] How to evaluate the quality of real-time voice dialogue between users and AI applications is a problem that remains to be solved. Summary of the Invention
[0008] This application provides an information processing method that can determine the interaction quality of a user engaging in real-time voice dialogue with an AI application.
[0009] Firstly, this application provides an information processing method applied to network devices through which uplink and downlink traffic for real-time voice dialogue between a user and an AI application passes. The method includes inputting the traffic characteristics of the uplink traffic and the traffic features of the downlink traffic into an AI model to determine the interaction latency of the real-time voice dialogue between the user and the AI application. The uplink traffic includes the audio stream sent by the user to the AI application via a terminal device, and the downlink traffic includes the audio stream sent by the AI application to the terminal device. Since the interaction latency of the real-time voice dialogue between the user and the AI application can characterize the interaction quality, for example, the lower the interaction latency, the better the interaction quality, and the higher the interaction latency, the worse the interaction quality. Therefore, after determining the interaction latency, the interaction quality of the real-time voice dialogue between the user and the AI application is further determined based on the interaction latency. Thus, this solution enables the evaluation of the interaction quality of real-time voice dialogue between a user and an AI application.
[0010] In one possible implementation, after inputting the uplink traffic characteristics into the AI model, key voice data in the uplink traffic can be determined based on the AI model's identification results. Similarly, by inputting the downlink traffic characteristics into the AI model, key voice data in the downlink traffic can be determined based on the AI model's identification results. After determining the key voice data in both the uplink and downlink traffic, the interaction latency for real-time voice dialogue between the user and the AI application can be determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic. In this solution, an AI model is used to assist in identifying key voice data in both the uplink and downlink traffic, thereby determining the interaction latency for real-time voice dialogue between the user and the AI application, so as to determine the interaction quality of real-time voice dialogue between the user and the AI application based on the interaction latency.
[0011] In one possible implementation, traffic characteristics are used to determine whether a message carries voice data. Therefore, the traffic characteristics can be features capable of distinguishing between messages carrying voice data and messages not carrying voice data. Alternatively, the traffic characteristics can be features capable of distinguishing the presence or absence of voice data. In one example, considering message size, messages carrying voice data can be distinguished from those not carrying voice data. Therefore, the traffic characteristics can include message size. In another example, considering packet transmission rate, messages carrying voice data can be distinguished from those not carrying voice data. Therefore, the traffic characteristics can include packet transmission rate. In yet another example, considering traffic burst characteristics, traffic burst characteristics can be distinguished from those containing voice data. Therefore, the traffic characteristics can include traffic burst characteristics.
[0012] In one possible implementation, considering that multiple real-time voice conversations may occur within a certain time period between a user and an AI application, the interaction latency of the real-time voice conversation between the user and the AI application can be the latency of any single real-time voice conversation among these multiple conversations. That is, the interaction latency is determined separately for each real-time voice conversation, thereby determining an interaction quality for each conversation separately. In another example, the interaction latency of the real-time voice conversation between the user and the AI application can be the average latency of multiple real-time voice conversations. That is, the average latency of multiple real-time voice conversations is used to determine the interaction quality for each conversation. The implementation method for determining the interaction quality is flexible.
[0013] In one possible implementation, the interaction latency and a latency threshold can be compared to determine the interaction quality. The latency threshold may include one or more thresholds. As a specific example, the latency threshold may include a single threshold; when the interaction latency is greater than or equal to this threshold, the interaction quality is poor; when the interaction latency is less than this threshold, the interaction quality is excellent. As another specific example, the latency threshold may include a first threshold and a second threshold; if the first threshold is less than the second threshold, then when the interaction latency is less than or equal to the first threshold, the interaction quality is excellent; when the interaction latency is greater than the first threshold and less than or equal to the second threshold, the interaction quality is average; and when the interaction quality is greater than the second threshold, the interaction quality is poor.
[0014] In one possible implementation, after determining the interaction quality, it can also be displayed. Specifically, the user can trigger the network device to display the interaction quality by entering a command via the command line. In this scenario, the network device can also be configured with a command to display the interaction quality, so that when the user enters the command to display the interaction quality via the command line, the network device can correctly respond to the command and thus display the interaction quality. Displaying the interaction quality refers to displaying information representing that interaction quality.
[0015] In one possible implementation, after determining the interaction quality, the network device can send indication information to the control and management entity. The indication information is used to indicate the interaction quality so that the control and management entity can perform corresponding processing measures based on the indication information. For example, after receiving the indication information, the control and management entity can display the interaction quality.
[0016] In one possible implementation, considering that in practical applications, if packet loss occurs in the audio data sent by the terminal device to the AI application, it may affect the AI application's correct understanding of the semantics of the user's voice data. Consequently, this will affect the AI application's inference efficiency, thereby reducing the efficiency of the AI application in returning corresponding responses to the user, and thus affecting the interaction quality of real-time voice dialogue between the user and the AI application. Therefore, the network device can also acquire first media data sent by the terminal device for real-time voice dialogue between the user and the AI application, the first media data including first audio data. Furthermore, when the interaction quality of the real-time voice dialogue does not meet the requirements, the network device sends the first media data and target data to the AI application, the target data being used for packet loss protection of the first media data. This approach ensures that the AI application receives the complete first media data, thereby guaranteeing the interaction quality of real-time voice dialogue between the user and the AI application.
[0017] In one possible implementation, the target data may be the first media data. In this scenario, after the network device acquires the first media data but before sending the first media data and the target data to the AI application, it may also copy the first media data to obtain multiple sets of first media data.
[0018] In one possible implementation, the target data may be verification information obtained by FEC encoding the first media data. In this scenario, after the network device acquires the first media data but before sending the first media data and the target data to the AI application, it can also perform FEC encoding on the first media data to obtain the target data.
[0019] In one possible implementation, the first audio data is key audio data for real-time voice dialogue between the user and the AI application. Alternatively, the first audio data is key audio data within the uplink traffic sent by the terminal device to the AI application. In this scenario, the first audio data can be all of the user's voice data or a portion of it. When the first audio data is a portion of the user's voice data, this portion can reflect the complete semantics of the user's entire voice data. In this case, network device resources can be used efficiently while ensuring that the AI application can correctly understand the semantics of the user's voice data and maintaining the quality of real-time voice dialogue between the user and the AI application.
[0020] In one possible implementation, where the first audio data is key voice data for a real-time voice dialogue between a user and an AI application, the network device can receive second audio data sent by the terminal device, representing the real-time voice dialogue between the user and the AI application. The second audio data includes non-voice data and the first audio data. Further, the first audio data within the second audio data is determined to extract it, thereby providing packet loss protection for the first audio data.
[0021] In one possible implementation, the determination of the first audio data in the second audio data can be achieved by inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0022] In one possible implementation, considering that when a user and an AI application engage in real-time voice dialogue, the data sent by the terminal device to the AI application may include not only audio data but also video data. Therefore, the first media data includes not only the first audio data but also the video data corresponding to the first audio data.
[0023] In one possible implementation, in a scenario where the first audio data is key audio data of a real-time voice dialogue between a user and an AI application, the first video data is a valid video corresponding to the key audio data of the real-time voice dialogue between the user and the AI application.
[0024] In one possible implementation, in a scenario where the first video data is valid video, the network device can receive second media data sent by the terminal device, representing a real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data representing the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. Further, the key voice data in the second audio data is determined to obtain the first audio data. After determining the first audio data, the first video data in the second video data can be determined based on the first audio data. After determining the first audio data and the first video data, packet loss protection can be applied to the first media data including the first audio data and the first video data.
[0025] Secondly, this application provides a method for training an AI model. The method includes: acquiring training samples, the training samples including: traffic features of a training audio stream and a label for each packet in the training audio stream, the label of each packet indicating whether each packet includes voice data. Then, based on the training samples, an AI model is trained, the AI model being used to identify whether a packet includes voice data. In a scenario where the interaction quality of real-time voice dialogue between a user and an AI application is evaluated, the AI model is used to identify whether packets included in uplink traffic include voice data, and to identify whether packets included in downlink traffic include voice data, so as to determine the interaction latency of real-time voice dialogue between the user and the AI application, thereby determining the interaction quality of real-time voice dialogue between the user and the AI application based on the interaction latency.
[0026] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0027] Thirdly, this application provides an information processing apparatus, the apparatus including a processing unit, the processing unit being configured to: input the traffic characteristics of uplink traffic and the traffic characteristics of downlink traffic into an artificial intelligence (AI) model respectively, to determine the interaction latency of real-time voice dialogue between a user and an AI application, wherein the uplink traffic includes: an audio stream sent by the user to the AI application through a terminal device, and the downlink traffic includes: an audio stream sent by the AI application to the terminal device; and determine the interaction quality of real-time voice dialogue between the user and the AI application based on the interaction latency.
[0028] In one possible implementation, the step of inputting the uplink traffic characteristics and downlink traffic characteristics into an artificial intelligence (AI) model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: inputting the uplink traffic characteristics into the AI model to determine key voice data in the uplink traffic; inputting the downlink traffic characteristics into the AI model to determine key voice data in the downlink traffic; and determining the interaction latency based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0029] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0030] In one possible implementation, the user and the AI application engage in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the interaction latency is the average latency of the multiple real-time voice conversations.
[0031] In one possible implementation, determining the interaction quality of the real-time voice dialogue between the user and the AI application based on the interaction latency includes: comparing the interaction latency with a latency threshold to determine the interaction quality.
[0032] In one possible implementation, the processing unit is further configured to: configure a command for displaying the interaction quality.
[0033] In one possible implementation, the apparatus further includes a sending unit for sending indication information to a control management entity, the indication information being used to indicate the interaction quality.
[0034] In one possible implementation, the apparatus further includes: an acquisition unit, configured to acquire first media data sent by the terminal device for real-time voice dialogue between the user and the AI application, the first media data including first audio data; the sending unit included in the apparatus is further configured to send the first media data and target data to the AI application when the interaction quality of the real-time voice dialogue does not meet the requirements, the target data being used for packet loss protection of the first media data.
[0035] In one possible implementation, the processing unit is further configured to: perform forward error correction (FEC) encoding on the first media data to obtain the target data before sending the first media data and the target data to the AI application; or, copy the first media data to obtain the target data before sending the first media data and the target data to the AI application.
[0036] In one possible implementation, the first audio data is key audio data from a real-time voice conversation between the user and the AI application.
[0037] In one possible implementation, the apparatus further includes: a receiving unit, configured to receive second audio data sent by the terminal device, representing a real-time voice dialogue between the user and the AI application, the second audio data including non-voice data and the first audio data; and a processing unit, further configured to determine the first audio data within the second audio data.
[0038] In one possible implementation, determining the first audio data in the second audio data includes: inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0039] In one possible implementation, the first media data further includes first video data corresponding to the first audio data.
[0040] In one possible implementation, the first video data is a valid video corresponding to key voice data of a real-time voice conversation between the user and the AI application.
[0041] In one possible implementation, the apparatus further includes: a receiving unit, configured to receive second media data sent by the terminal device for real-time voice dialogue between the user and the AI application, the second media data including second audio data and second video data, the second audio data including non-voice data and key voice data of the voice dialogue between the user and the AI application, and the second video data including invalid video and the valid video; the processing unit is further configured to determine the key voice data in the second audio data to obtain the first audio data; and to determine the first video data in the second video data based on the first audio data.
[0042] Fourthly, this application provides an AI model training apparatus, the apparatus comprising: an acquisition unit for acquiring training samples, the training samples including: traffic characteristics of a training audio stream and a label for each message in the training audio stream, the label for each message indicating whether each message includes voice data; and a processing unit for training an AI model based on the training samples, the AI model being used to identify whether a message includes voice data.
[0043] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0044] This application also provides a communication method and apparatus that can ensure the quality of real-time voice dialogue between users and AI applications.
[0045] Fifthly, this application provides a communication method applied to a network device. The method includes: acquiring first media data sent by a terminal device for a real-time voice dialogue between a user and an AI application, wherein the first media data includes first audio data. To avoid packet loss during transmission of the first audio data, which could affect the AI model's understanding of the user's voice data, after acquiring the first media data, it can be copied to obtain multiple sets of first media data, which are then sent to the AI application. In this application, since multiple sets of first media data are sent to the AI application, if a portion of a message in one set of first media data is lost, it can be recovered using corresponding messages from other sets of first media data. This allows the AI application to receive the complete first media data, which in turn helps the AI application correctly understand the user's voice data, thereby improving the AI application's inference efficiency and ensuring the quality of real-time voice dialogue between the user and the AI application.
[0046] In one possible implementation, the first audio data is key audio data for real-time voice dialogue between the user and the AI application. Alternatively, the first audio data is key audio data within the uplink traffic sent by the terminal device to the AI application. In this scenario, the first audio data can be all of the user's voice data or a portion of it. When the first audio data is a portion of the user's voice data, this portion can reflect the complete semantics of the user's entire voice data. In this case, network device resources can be used efficiently while ensuring that the AI application can correctly understand the semantics of the user's voice data and maintaining the quality of real-time voice dialogue between the user and the AI application.
[0047] In one possible implementation, where the first audio data is key voice data for a real-time voice dialogue between a user and an AI application, the network device can receive second audio data sent by the terminal device, representing the real-time voice dialogue between the user and the AI application. The second audio data includes non-voice data and the first audio data. Further, the first audio data within the second audio data is determined to extract it, thereby providing packet loss protection for the first audio data.
[0048] In one possible implementation, the determination of the first audio data in the second audio data can be achieved by inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0049] In one possible implementation, considering that when a user and an AI application engage in real-time voice dialogue, the data sent by the terminal device to the AI application may include not only audio data but also video data. Therefore, the first media data includes not only the first audio data but also the video data corresponding to the first audio data.
[0050] In one possible implementation, in a scenario where the first audio data is key audio data of a real-time voice dialogue between a user and an AI application, the first video data is a valid video corresponding to the key audio data of the real-time voice dialogue between the user and the AI application.
[0051] In one possible implementation, in a scenario where the first video data is valid video, the network device can receive second media data sent by the terminal device, representing a real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data representing the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. Further, the key voice data in the second audio data is determined to obtain the first audio data. After determining the first audio data, the first video data in the second video data can be determined based on the first audio data. After determining the first audio data and the first video data, the first media data, including the first audio data and the first video data, can be copied to achieve packet loss protection for the first media data.
[0052] In one possible implementation, traffic characteristics are used to determine whether a message carries voice data. Therefore, the traffic characteristics can be features capable of distinguishing between messages carrying voice data and messages not carrying voice data. Alternatively, the traffic characteristics can be features capable of distinguishing the presence or absence of voice data. In one example, considering message size, messages carrying voice data can be distinguished from those not carrying voice data. Therefore, the traffic characteristics can include message size. In another example, considering packet transmission rate, messages carrying voice data can be distinguished from those not carrying voice data. Therefore, the traffic characteristics can include packet transmission rate. In yet another example, considering traffic burst characteristics, traffic burst characteristics can be distinguished from those containing voice data. Therefore, the traffic characteristics can include traffic burst characteristics.
[0053] In one possible implementation, considering that network quality issues may lead to packet loss in the first media data sent by the network device to the AI application, consequently affecting the interaction quality of real-time voice dialogue between the user and the AI application, the network device can, in one example, copy the first media data when network quality is unsatisfactory. This allows the receiving device on the application side to recover the lost packets based on multiple copies of the first media data.
[0054] In one possible implementation, considering that if the interaction quality of real-time voice dialogue between the user and the AI application does not meet the requirements, it may be due to packet loss in the uplink traffic sent by the user to the AI application. Therefore, when the interaction quality of real-time voice dialogue between the user and the AI application does not meet the requirements, the network device can copy the first media data, thereby enabling the receiving device on the application side to recover the lost packets based on multiple first media data sets, thus improving the interaction quality of real-time voice dialogue between the user and the AI application.
[0055] In one possible implementation, if both network quality and the interaction quality of real-time voice dialogue fail to meet requirements, it indicates that low network quality may be causing packet loss in the uplink traffic sent by the user to the AI application, further resulting in the failure to meet the interaction quality requirements for real-time voice dialogue between the user and the AI application. Therefore, when both interaction quality and network quality fail to meet requirements, the network device can copy the first media data, enabling the receiving device on the application side to recover the lost packets based on multiple sets of first media data, thereby improving the interaction quality of real-time voice dialogue between the user and the AI application.
[0056] In one possible implementation, the network device can input the traffic characteristics of uplink traffic and downlink traffic into the AI model to determine the interaction latency of real-time voice dialogue between the user and the AI application. Uplink traffic includes the audio stream sent by the user to the AI application through the terminal device, and downlink traffic includes the audio stream sent by the AI application to the terminal device. Since the interaction latency of real-time voice dialogue between the user and the AI application can characterize the interaction quality, for example, the lower the interaction latency, the better the interaction quality, and the higher the interaction latency, the worse the interaction quality, after determining the interaction latency, the interaction quality of real-time voice dialogue between the user and the AI application can be further determined based on the interaction latency.
[0057] In one possible implementation, after inputting the uplink traffic characteristics into the AI model, key voice data in the uplink traffic can be determined based on the AI model's identification results. Similarly, by inputting the downlink traffic characteristics into the AI model, key voice data in the downlink traffic can be determined based on the AI model's identification results. After determining the key voice data in both the uplink and downlink traffic, the interaction latency for real-time voice dialogue between the user and the AI application can be determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic. In this solution, an AI model is used to assist in identifying key voice data in both the uplink and downlink traffic, thereby determining the interaction latency for real-time voice dialogue between the user and the AI application, so as to determine the interaction quality of real-time voice dialogue between the user and the AI application based on the interaction latency.
[0058] In one possible implementation, considering that multiple real-time voice conversations may occur within a certain time period between a user and an AI application, the interaction latency of the real-time voice conversation between the user and the AI application can be the latency of any single real-time voice conversation among these multiple conversations. That is, the interaction latency is determined separately for each real-time voice conversation, thereby determining an interaction quality for each conversation separately. In another example, the interaction latency of the real-time voice conversation between the user and the AI application can be the average latency of multiple real-time voice conversations. That is, the average latency of multiple real-time voice conversations is used to determine the interaction quality for each conversation. The implementation method for determining the interaction quality is flexible.
[0059] In one possible implementation, the interaction latency and a latency threshold can be compared to determine the interaction quality. The latency threshold may include one or more thresholds. As a specific example, the latency threshold may include a single threshold; when the interaction latency is greater than or equal to this threshold, the interaction quality is poor; when the interaction latency is less than this threshold, the interaction quality is excellent. As another specific example, the latency threshold may include a first threshold and a second threshold; if the first threshold is less than the second threshold, then when the interaction latency is less than or equal to the first threshold, the interaction quality is excellent; when the interaction latency is greater than the first threshold and less than or equal to the second threshold, the interaction quality is average; and when the interaction quality is greater than the second threshold, the interaction quality is poor.
[0060] In one possible implementation, after determining the interaction quality, it can also be displayed. Specifically, the user can trigger the network device to display the interaction quality by entering a command via the command line. In this scenario, the network device can also be configured with a command to display the interaction quality, so that when the user enters the command to display the interaction quality via the command line, the network device can correctly respond to the command and thus display the interaction quality. Displaying the interaction quality refers to displaying information representing that interaction quality.
[0061] In one possible implementation, after determining the interaction quality, the network device can send indication information to the control and management entity. The indication information is used to indicate the interaction quality so that the control and management entity can perform corresponding processing measures based on the indication information. For example, after receiving the indication information, the control and management entity can display the interaction quality.
[0062] Sixthly, this application provides a communication device, the device comprising: a processing unit, configured to acquire first media data sent by a terminal device for a real-time voice dialogue between a user and an artificial intelligence (AI) application, the first media data including first audio data; copy the first media data to obtain a plurality of first media data; and a sending unit, configured to send the plurality of first media data to the AI application.
[0063] In one possible implementation, the first audio data is key audio data from a real-time voice conversation between the user and the AI application.
[0064] In one possible implementation, the apparatus further includes: a receiving unit for receiving second audio data sent by the terminal device, the second audio data including non-voice data and the first audio data; and a processing unit for determining the first audio data in the second audio data.
[0065] In one possible implementation, determining the first audio data in the second audio data includes: inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0066] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0067] In one possible implementation, the first media data further includes first video data corresponding to the first audio data.
[0068] In one possible implementation, the first video data is a valid video corresponding to key voice data of a real-time voice conversation between the user and the AI application.
[0069] In one possible implementation, the apparatus further includes: a receiving unit, configured to receive second media data sent by the terminal device, the second media data including second audio data and second video data, the second audio data including non-voice data and key voice data of a voice dialogue between the user and the AI application, and the second video data including invalid video and the valid video; the processing unit is further configured to determine the key voice data in the second audio data to obtain the first audio data; and to determine the first video data in the second video data based on the first audio data.
[0070] In one possible implementation, copying the first media data includes: copying the first media data when the network quality and / or the interaction quality of the real-time voice dialogue do not meet the requirements.
[0071] In one possible implementation, the processing unit is further configured to: input the uplink traffic characteristics and downlink traffic characteristics into the AI model respectively to determine the interaction latency of the real-time voice dialogue between the user and the AI application, wherein the uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device; and determine the interaction quality based on the interaction latency.
[0072] In one possible implementation, the step of inputting the uplink traffic characteristics and downlink traffic characteristics into an AI speech recognition model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: inputting the uplink traffic characteristics into the AI model to determine key voice data in the uplink traffic; inputting the downlink traffic characteristics into the AI model to determine key voice data in the downlink traffic; and determining the interaction latency based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0073] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0074] In one possible implementation, the user and the AI application engage in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the interaction latency is the average latency of the multiple real-time voice conversations.
[0075] In one possible implementation, determining the interaction quality based on the interaction latency includes comparing the interaction latency with a latency threshold to determine the interaction quality.
[0076] This application provides an alternative communication method and device that can ensure the quality of real-time voice dialogue between users and AI applications.
[0077] Seventhly, this application provides a communication method applied to a target network device. The method includes: receiving first media data for a real-time voice dialogue between a user and an AI application, the first media data including first audio data. When network quality and / or the interaction quality of the real-time voice dialogue does not meet requirements, packet loss protection is initiated for the first media data. Specifically, the first media data and target data are sent to a receiving device, the target data being used to perform packet loss protection on the first media data. In this application, after receiving the first media data, the network device, in addition to sending the first media data to the receiving device, also sends the target data for packet loss protection of the first media data to the receiving device, thereby enabling the receiving device to receive the complete first media data, thereby improving the interaction quality of the real-time voice dialogue between the user and the AI application.
[0078] In one possible implementation, the target data may be the first media data. In this scenario, after the network device receives the first media data but before sending the first media data and the target data to the receiving device, it may also copy the first media data to obtain multiple sets of first media data.
[0079] In one possible implementation, the target data may be verification information obtained by FEC encoding the first media data. In this scenario, after the network device receives the first media data but before sending the first media data and the target data to the receiving device, it may also perform FEC encoding on the first media data to obtain the target data.
[0080] In one possible implementation, the first media data is sent by a terminal device; in other words, the target network device can receive the first media data sent by the terminal device. In this scenario, the receiving device is an AI application, meaning the target network device can send the first media data and the target data to the AI application. In this scenario, packet loss protection can be implemented for the first media data sent from the terminal device to the AI application, thereby ensuring the quality of real-time voice dialogue between the user and the AI application.
[0081] In one possible implementation, the first media data is sent by the AI application; in other words, the target network device can receive the first media data sent by the AI application. In this scenario, the receiving device is a terminal device, meaning the target network device can send the first media data and the target data to the terminal device. In this scenario, packet loss protection can be implemented for the first media data sent by the AI application to the terminal device, thereby ensuring the quality of real-time voice dialogue between the user and the AI application.
[0082] In one possible implementation, the target network device can input the traffic characteristics of uplink traffic and downlink traffic into an AI model to determine the interaction latency of real-time voice dialogue between the user and the AI application. Uplink traffic includes the audio stream sent by the user to the AI application via the terminal device, and downlink traffic includes the audio stream sent by the AI application to the terminal device. Since the interaction latency of real-time voice dialogue between the user and the AI application can characterize the interaction quality, for example, the lower the interaction latency, the better the interaction quality, and the higher the interaction latency, the worse the interaction quality, after determining the interaction latency, the interaction quality of real-time voice dialogue between the user and the AI application can be further determined based on the interaction latency.
[0083] In one possible implementation, after inputting the uplink traffic characteristics into the AI model, key voice data in the uplink traffic can be determined based on the AI model's identification results. Similarly, by inputting the downlink traffic characteristics into the AI model, key voice data in the downlink traffic can be determined based on the AI model's identification results. After determining the key voice data in both the uplink and downlink traffic, the interaction latency for real-time voice dialogue between the user and the AI application can be determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic. In this solution, an AI model is used to assist in identifying key voice data in both the uplink and downlink traffic, thereby determining the interaction latency for real-time voice dialogue between the user and the AI application, so as to determine the interaction quality of real-time voice dialogue between the user and the AI application based on the interaction latency.
[0084] In one possible implementation, considering that multiple real-time voice conversations may occur within a certain time period between a user and an AI application, the interaction latency of the real-time voice conversation between the user and the AI application can be the latency of any single real-time voice conversation among these multiple conversations. That is, the interaction latency is determined separately for each real-time voice conversation, thereby determining an interaction quality for each conversation separately. In another example, the interaction latency of the real-time voice conversation between the user and the AI application can be the average latency of multiple real-time voice conversations. That is, the average latency of multiple real-time voice conversations is used to determine the interaction quality for each conversation. The implementation method for determining the interaction quality is flexible.
[0085] In one possible implementation, the interaction latency and a latency threshold can be compared to determine the interaction quality. The latency threshold may include one or more thresholds. As a specific example, the latency threshold may include a single threshold; when the interaction latency is greater than or equal to this threshold, the interaction quality is poor; when the interaction latency is less than this threshold, the interaction quality is excellent. As another specific example, the latency threshold may include a first threshold and a second threshold; if the first threshold is less than the second threshold, then when the interaction latency is less than or equal to the first threshold, the interaction quality is excellent; when the interaction latency is greater than the first threshold and less than or equal to the second threshold, the interaction quality is average; and when the interaction quality is greater than the second threshold, the interaction quality is poor.
[0086] In one possible implementation, after determining the interaction quality, it can also be displayed. Specifically, the user can trigger the target network device to display the interaction quality by entering a command via a command line. In this scenario, the target network device can also be configured with a command for displaying the interaction quality, so that when the user enters the command to display the interaction quality via the command line, the target network device can correctly respond to the command and thus display the interaction quality. Displaying the interaction quality refers to displaying information representing that interaction quality.
[0087] In one possible implementation, after determining the interaction quality, the target network device can send indication information to the control and management entity. The indication information is used to indicate the interaction quality so that the control and management entity can perform corresponding processing measures based on the indication information. For example, after receiving the indication information, the control and management entity can display the interaction quality.
[0088] In one possible implementation, traffic characteristics are used to determine whether a message carries voice data. Therefore, the traffic characteristics can be features capable of distinguishing between messages carrying voice data and messages not carrying voice data. Alternatively, the traffic characteristics can be features capable of distinguishing the presence or absence of voice data. In one example, considering message size, messages carrying voice data can be distinguished from those not carrying voice data. Therefore, the traffic characteristics can include message size. In another example, considering packet transmission rate, messages carrying voice data can be distinguished from those not carrying voice data. Therefore, the traffic characteristics can include packet transmission rate. In yet another example, considering traffic burst characteristics, traffic burst characteristics can be distinguished from those containing voice data. Therefore, the traffic characteristics can include traffic burst characteristics.
[0089] In one possible implementation, the first audio data is key audio data for real-time voice dialogue between the user and the AI application. Specifically, when the first audio data is audio data sent from the terminal device to the AI application, the key audio data for real-time voice dialogue between the user and the AI application can be all of the user's voice data or a portion of the user's voice data; this embodiment does not impose a specific limitation. When the key audio data is a portion of the user's voice data, this portion of voice data can reflect the complete semantics of the user's entire voice data. Similarly, when the first audio data is audio data sent from the AI application to the terminal device, the key audio data for real-time voice dialogue between the user and the AI application can be all of the AI application's voice data or a portion of the AI application's voice data; this embodiment does not impose a specific limitation. When the key audio data is a portion of the AI application's voice data, this portion of voice data can reflect the complete semantics of the AI application's entire voice data. In this case, network device resources can be reasonably utilized while ensuring the interaction quality of real-time voice dialogue between the user and the AI application.
[0090] In one possible implementation, where the first audio data is key voice data for a real-time voice dialogue between a user and an AI application, the network device can receive second audio data for the real-time voice dialogue between the user and the AI application. The second audio data includes non-voice data and the first audio data. Further, the first audio data is determined from the second audio data to extract it, thereby providing packet loss protection for the first audio data.
[0091] In one possible implementation, the determination of the first audio data in the second audio data can be achieved by inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0092] In one possible implementation, considering that when a user and an AI application engage in real-time voice dialogue, the data exchanged between the terminal device and the AI application may include not only audio data but also video data. Therefore, the first media data includes not only the first audio data but also the video data corresponding to the first audio data.
[0093] In one possible implementation, in a scenario where the first audio data is key audio data of a real-time voice dialogue between a user and an AI application, the first video data is a valid video corresponding to the key audio data of the real-time voice dialogue between the user and the AI application.
[0094] In one possible implementation, in a scenario where the first video data is valid video, the network device can receive second media data of real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data of the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. Further, the key voice data in the second audio data is determined to obtain the first audio data. After determining the first audio data, the first video data in the second video data can be determined based on the first audio data. After determining the first audio data and the first video data, packet loss protection can be applied to the first media data including the first audio data and the first video data.
[0095] Eighthly, this application provides a communication device, the device comprising: a receiving unit for receiving first media data for real-time voice dialogue between a user and an artificial intelligence (AI) application, the first media data including first audio data; and a sending unit for sending the first media data and target data to a receiving device when the network quality and / or the interaction quality of the real-time voice dialogue does not meet the requirements, the target data being used to protect the first media data from packet loss.
[0096] In one possible implementation, the apparatus further includes a processing unit configured to: perform forward error correction (FEC) encoding on the first media data to obtain the target data before sending the first media data and the target data to the receiving device; or, copy the first media data to obtain the target data before sending the first media data and the target data to the receiving device.
[0097] In one possible implementation, the receiving unit is configured to: receive the first media data sent by the terminal device; and the sending unit is configured to: send the first media data and the target data to the AI application.
[0098] In one possible implementation, the receiving unit is configured to: receive the first media data sent by the AI application; and the sending unit is configured to: send the first media data and the target data to the terminal device.
[0099] In one possible implementation, the processing unit of the device is further configured to: input the traffic characteristics of uplink traffic and the traffic characteristics of downlink traffic into the AI model respectively, to determine the interaction latency of real-time voice dialogue between the user and the AI application, wherein the uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device; and determine the interaction quality based on the interaction latency.
[0100] In one possible implementation, the step of inputting the uplink traffic characteristics and downlink traffic characteristics into an AI speech recognition model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: inputting the uplink traffic characteristics into the AI model to determine key voice data in the uplink traffic; inputting the downlink traffic characteristics into the AI model to determine key voice data in the downlink traffic; and determining the interaction latency based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0101] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0102] In one possible implementation, the user and the AI application engage in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the interaction latency is the average latency of the multiple real-time voice conversations.
[0103] In one possible implementation, determining the interaction quality based on the interaction latency includes comparing the interaction latency with a latency threshold to determine the interaction quality.
[0104] In one possible implementation, the first audio data is key audio data from a real-time voice conversation between the user and the AI application.
[0105] In one possible implementation, the receiving unit is further configured to receive second audio data of a real-time voice conversation between the user and the AI application, the second audio data including non-voice data and the first audio data; the processing unit is further configured to determine the first audio data in the second audio data.
[0106] In one possible implementation, the processing unit is configured to: input the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0107] In one possible implementation, the first media data further includes first video data corresponding to the first audio data.
[0108] In one possible implementation, the first video data is a valid video corresponding to key voice data of a real-time voice conversation between the user and the AI application.
[0109] In one possible implementation, the receiving unit is further configured to receive second media data for real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data for the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. The processing unit is further configured to determine the key voice data in the second audio data to obtain the first audio data. Based on the first audio data, the processing unit determines the first video data in the second video data.
[0110] Ninthly, embodiments of this application provide a communication device, including: a processor and a memory;
[0111] The memory is used to store instructions; the processor is used to execute the instructions, causing the communication device to perform the method described in the first aspect and any one of the first aspects above, or causing the communication device to perform the method described in the second aspect and any one of the second aspects above, or causing the communication device to perform the method described in the fifth aspect and any one of the fifth aspects above, or causing the communication device to perform the method described in the seventh aspect and any one of the seventh aspects above.
[0112] In a tenth aspect, embodiments of this application provide a communication device, including a processor, the processor being configured to perform operations other than transmission and reception operations in any of the methods described in the first aspect and above; or, the processor being configured to perform operations other than transmission and reception operations in any of the methods described in the second aspect and above; or, the processor being configured to perform operations other than transmission and reception operations in any of the methods described in the fifth aspect and above; or, the processor being configured to perform operations other than transmission and reception operations in any of the methods described in the seventh aspect and above.
[0113] In one possible implementation, the communication device further includes a communication interface connected to the processor. The communication interface is used to perform the transmit / receive operations described in the first aspect and any one of the methods described in the first aspect above; or, the communication interface is used to perform the transmit / receive operations described in the second aspect and any one of the methods described in the second aspect above; or, the communication interface is used to perform the transmit / receive operations described in the fifth aspect and any one of the methods described in the fifth aspect above; or, the communication interface is used to perform the transmit / receive operations described in the seventh aspect and any one of the methods described in the seventh aspect above.
[0114] Eleventhly, embodiments of this application provide a computer-readable storage medium including instructions or a computer program that, when executed on a processor, implements the method described in the first aspect and any one of the first aspects above; or, when executed on a processor, implements the method described in the second aspect and any one of the second aspects above; or, when executed on a processor, implements the method described in the fifth aspect and any one of the fifth aspects above; or, when executed on a processor, implements the method described in the seventh aspect and any one of the seventh aspects above.
[0115] In a twelfth aspect, embodiments of this application provide a computer program product, including a computer program product that, when running on a processor, implements the method described in the first aspect and any one of the first aspects above, or, when running on a processor, implements the method described in the second aspect and any one of the second aspects above, or, when running on a processor, implements the method described in the fifth aspect and any one of the fifth aspects above, or, when running on a processor, implements the method described in the seventh aspect and any one of the seventh aspects above. Attached Figure Description
[0116] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0117] Figure 1a is a schematic diagram of an application scenario provided by an embodiment of this application;
[0118] Figure 1b is a schematic diagram of another application scenario provided by an embodiment of this application;
[0119] Figure 2 is a flowchart illustrating an information processing method provided in this application;
[0120] Figure 3 is a flowchart illustrating a communication method provided in an embodiment of this application;
[0121] Figure 4a is a schematic diagram of an exemplary application scenario provided by an embodiment of this application;
[0122] Figure 4b is a schematic diagram of another exemplary application scenario provided by the embodiments of this application;
[0123] Figure 5a is a flowchart illustrating a training method for an AI model provided in an embodiment of this application;
[0124] Figure 5b is a flowchart illustrating a communication method provided in this application;
[0125] Figure 5c is a flowchart illustrating another communication method provided in an embodiment of this application;
[0126] Figure 6 is a schematic diagram of a communication method provided in an embodiment of this application;
[0127] Figure 7 is a schematic diagram of a communication method provided in an embodiment of this application;
[0128] Figure 8a is a schematic diagram of the structure of an information processing device provided in an embodiment of this application;
[0129] Figure 8b is a schematic diagram of the structure of a communication device provided in an embodiment of this application;
[0130] Figure 9 is a schematic diagram of the structure of another AI model training device provided in an embodiment of this application;
[0131] Figure 10 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;
[0132] Figure 11 is a schematic diagram of the structure of a communication device provided in an embodiment of this application. Detailed Implementation
[0133] This application provides an information processing method that can determine the interaction quality of a user engaging in real-time voice dialogue with an AI application.
[0134] Users can interact with AI applications through terminal devices, such as those deploying a large language model (LLM). The LLM mentioned here at least has the ability to process audio, and optionally, it also has the ability to process video.
[0135] In one example, a user can send voice data to an AI application via a terminal device. The AI application then understands the voice data from the terminal device, infers a corresponding response, and replies to the user via voice. Specifically, the AI application can send its voice response to the user back to the terminal device. In another example, the data sent by the user to the AI application via the terminal device includes both voice and video data. The AI application can then understand both the voice and video data from the terminal device and infer a corresponding response. Furthermore, the AI application can respond to the user using both voice and video.
[0136] Please refer to Figure 1a for understanding. Figure 1a is a schematic diagram of an application scenario provided by an embodiment of this application.
[0137] As shown in Figure 1a, the terminal device sends voice data to the AI application through network device 1, and correspondingly, the terminal device receives voice data sent by the AI application through network device 1. Wherein:
[0138] Network device 1 is a network device that interfaces with the terminal device. For example, network device 1 can be an access-side device, such as an access router (AR) or an access point (AP). In this scenario, although not shown in Figure 1a, the voice data sent by the terminal device to the AI application through network device 1 can be processed by the receiving device on the application side and then sent to the AI application. The receiving device on the AI application side mentioned here can be a server, and this application does not limit it.
[0139] In one example, refer to Figure 1b, which is a schematic diagram of another application scenario provided by an embodiment of this application. Voice data received by network device 1 from the terminal device is sent to the AI application via network device 2. Similarly, network device 2 receives voice data sent by the AI application and sends the voice data back to network device 1, which then sends the voice data to the terminal device.
[0140] In one example, the AI application shown in Figure 1a can be deployed in the cloud.
[0141] In one example, the AI application and network device 2 shown in Figure 1b can be deployed, for example, in a data center.
[0142] Currently, for network devices, such as network device 1 shown in Figures 1a and 1b, how to evaluate the interaction quality of real-time voice dialogue between users and AI applications is an unsolved problem.
[0143] To address the aforementioned problems, this application provides an information processing method. The information processing method provided by this application will now be described in conjunction with the accompanying drawings.
[0144] See Figure 2, which is a flowchart illustrating an information processing method provided in this application.
[0145] The method 100 shown in Figure 2 can be applied to the scenario shown in Figure 1a or Figure 1b, wherein the method 100 can be executed by the network device 1.
[0146] Method 100 includes the following steps S101-S102.
[0147] S101: Input the uplink traffic characteristics and downlink traffic characteristics into the AI model respectively to determine the interaction latency of real-time voice dialogue between the user and the AI application. The uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device.
[0148] As described in the previous scenario of Figure 1a, both uplink and downlink traffic need to pass through network device 1 during transmission. Therefore, network device 1 can acquire both uplink and downlink traffic.
[0149] Uplink traffic includes audio streams sent by users to AI applications via their terminal devices. In this application, the audio streams sent by users to AI applications via their terminal devices include at least voice data, i.e., the user's voice data. Additionally, the audio streams sent by users to AI applications via their terminal devices may also include non-voice data, such as ambient noise.
[0150] Downlink traffic includes audio streams sent by the AI application to the terminal device. In this application, the audio stream sent by the AI application to the terminal device includes at least voice data, that is, voice data generated by the AI application. Similarly, the audio stream sent by the AI application to the terminal device may also include non-voice data.
[0151] After acquiring uplink and downlink traffic, network device 1 can determine the traffic characteristics of uplink traffic and the traffic features of downlink traffic, and input the traffic characteristics of uplink traffic and downlink traffic into the AI model respectively.
[0152] In this application, the network device 1 is equipped with the AI model, which is used to identify whether a packet includes voice data. Specifically, after inputting the traffic characteristics of uplink traffic into the AI model, the AI model can identify whether each packet in the uplink traffic includes voice data. After inputting the traffic characteristics of downlink traffic into the AI model, the AI model can identify whether each packet in the downlink traffic includes voice data.
[0153] After inputting the traffic characteristics of the uplink traffic into the AI model, network device 1 can determine the key voice data in the uplink traffic based on the results of the AI model's identification of the uplink traffic characteristics. The key voice data in the uplink traffic can be all of the user's voice data or a portion of the user's voice data; this embodiment does not impose a specific limitation. When the key voice data is a portion of the user's voice data, this portion of voice data can reflect the complete semantics of the user's total voice data.
[0154] In one example, the key voice data in the uplink traffic consists of multiple packets in the uplink traffic that include voice data.
[0155] For example:
[0156] The uplink traffic comprises 100 packets. The AI model identifies the following based on the traffic characteristics of the uplink traffic: packets 1 to 50 contain voice data, while packets 51 to 100 do not contain voice data. Therefore, the key voice data in the uplink traffic includes packets 1 to 50. Alternatively, the uplink traffic comprises 100 packets. The AI model identifies the following based on the traffic characteristics of the uplink traffic: packets 1 to 25 and packets 27 to 50 all contain voice data, while packets 26 and packets 51 to 100 do not contain voice data. Therefore, the key voice data in the uplink traffic includes packets 1 to 25 and packets 27 to 50.
[0157] In another example, when most of the consecutive packets in the uplink traffic include voice data, and only a few packets do not include voice data, the key voice data in the uplink traffic is the aforementioned consecutive packets.
[0158] For example:
[0159] The uplink traffic includes 100 packets. The AI model identifies the following based on the traffic characteristics of the uplink traffic: packets 1 to 25 and packets 27 to 50 all contain voice data, while packets 26 and packets 51 to 100 do not contain voice data. Therefore, the key voice data in the uplink traffic includes packets 1 to 50.
[0160] Similarly, after inputting the downlink traffic characteristics into the AI model, the key voice data in the downlink traffic can be determined based on the AI model's identification of the downlink traffic characteristics. Similar to the key voice data in uplink traffic, the key voice data in downlink traffic can be all the voice data of the AI application or only a portion of the voice data; this application does not impose specific limitations on this.
[0161] Similar to key voice data in uplink traffic, in one example, key voice data in downlink traffic refers to multiple packets in the downlink traffic that include voice data. In another example, when most packets in a series of consecutive packets in the downlink traffic include voice data, and only a few packets do not include voice data, the key voice data in the downlink traffic refers to the aforementioned series of consecutive packets in the downlink traffic.
[0162] After identifying the key voice data in the uplink traffic and the key voice data in the downlink traffic, the interaction latency between the user and the AI application for real-time voice dialogue can be determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0163] In this application, the uplink traffic includes Real-time Transport Protocol (RTP) messages. Each RTP message includes a timestamp field to carry the message's transmission timestamp. For uplink traffic, the timestamp in the message is the timestamp when the terminal device sends the message. In one example, the critical voice data ending time in the aforementioned uplink traffic can be the transmission time of the last message in the critical voice data of the uplink traffic, that is, the timestamp carried in the last message. After determining the critical voice data in the uplink traffic, network device 1 can extract the timestamp from the last message in the critical voice data of the uplink traffic to obtain the ending time of the critical voice data in the uplink traffic. In another example, the critical voice data ending time in the aforementioned uplink traffic can be the time when network device 1 receives the last message in the critical voice data of the uplink traffic. In this example, network device 1 records the time it receives each message in the uplink traffic. After determining the key voice data in the uplink traffic, network device 1 can obtain its own recorded time of receiving the last message, thereby obtaining the end time of the key voice data in the uplink traffic. Since network device 1 is a network device that interfaces with the terminal device, the difference between the sending time of the last message in the key voice data in the uplink traffic and the time when network device 1 receives the last message is very small.
[0164] In this application, the start time of the key voice data in the downlink traffic is the time when network device 1 receives the first packet of the key voice data in the downlink traffic. Specifically, network device 1 records the time it receives each packet in the downlink traffic. Therefore, after determining the key voice data in the downlink traffic, network device 1 can obtain its recorded time of receiving the first packet, thereby obtaining the start time of the key voice data in the downlink traffic.
[0165] After obtaining the end time and the start time, network device 1 can subtract the end time from the start time to obtain the interaction delay. This interaction delay is the difference between the time when the terminal device sends the last message carrying the user's voice data and the time when network device 1 receives the first message carrying the AI application's voice data. Since network device 1 is a network device that interfaces with the terminal device, the aforementioned start time is equivalent to the time when the terminal device receives the first message carrying the AI application's voice data. Therefore, the aforementioned difference can represent the delay between the user ending their question to the AI application and the user beginning to receive the AI application's response. Thus, this interaction delay can characterize the interaction quality of real-time voice dialogue between the user and the AI application.
[0166] As described above, traffic characteristics are used to determine whether a message carries voice data. Therefore, the traffic characteristics can be features that distinguish between messages carrying voice data and messages not carrying voice data. Alternatively, the traffic characteristics can be features that distinguish whether voice data is present.
[0167] In one example, considering that message size can distinguish between messages carrying voice data and messages not carrying voice data, specifically, messages carrying voice data are larger (i.e., longer), while messages not carrying voice data are smaller. Therefore, the traffic characteristics may include message size.
[0168] In another example, the packet sending rate can distinguish between packets carrying voice data and packets not carrying voice data. Here, packet sending rate refers to the rate of packets sent within a certain time period. Specifically, the packet sending rate is higher when voice data is present and lower when no voice data is present. In other words, a higher packet sending rate indicates a higher probability that the packet contains voice data, while a lower packet sending rate indicates a lower probability. Therefore, the traffic characteristic can include packet sending rate. The time period mentioned here is a relatively small period, for example, 10 milliseconds.
[0169] In another example, considering that burst traffic characteristics can distinguish the presence of voice data, for example, when a user suddenly starts speaking, the uplink traffic received by network device 1 will generally exhibit a burst traffic pattern, such as 20 packets bursting within 10 ms. Therefore, the traffic characteristics can include burst traffic features.
[0170] Considering that users and AI applications may engage in multiple real-time voice conversations within a certain time period, in one example, the interaction latency of a real-time voice conversation between the user and the AI application can be the latency of any single real-time voice conversation among these multiple conversations. That is, for any single real-time voice conversation, S101 can be used to determine the latency of that conversation. In another example, the interaction latency of a real-time voice conversation between the user and the AI application can be the average latency of multiple real-time voice conversations. The latency of any single real-time voice conversation can be determined by inputting the uplink traffic characteristics and downlink traffic characteristics corresponding to that conversation into the AI model. For details, please refer to the description of S101 in method 100 above; it will not be repeated here.
[0171] S102: Determine the interaction quality of the real-time voice dialogue between the user and the AI application based on the interaction latency.
[0172] After determining the interaction latency of the real-time voice dialogue between the user and the AI application, the interaction quality of the real-time voice dialogue between the user and the AI application can be determined based on the interaction latency.
[0173] In one example, the interaction latency can be directly used as the interaction quality.
[0174] In another example, the interaction latency and a latency threshold can be compared to determine the interaction quality. The latency threshold may include one or more thresholds. As a specific example, the latency threshold includes a single threshold; when the interaction latency is greater than or equal to this threshold, the interaction quality is poor; when the interaction latency is less than this threshold, the interaction quality is excellent. As another specific example, the latency threshold includes a first threshold and a second threshold; if the first threshold is less than the second threshold, then when the interaction latency is less than or equal to the first threshold, the interaction quality is excellent; when the interaction latency is greater than the first threshold and less than or equal to the second threshold, the interaction quality is average; and when the interaction quality is greater than the second threshold, the interaction quality is poor.
[0175] As described above, the solution of this application analyzes the traffic characteristics of uplink and downlink traffic using an AI model. Based on the results of the analysis of the traffic characteristics of uplink and downlink traffic using the AI model, the interaction latency of real-time voice dialogue between the user and the AI application is determined. Furthermore, based on the interaction latency, the interaction quality of real-time voice dialogue between the user and the AI application is determined, thus realizing the evaluation of the interaction quality of real-time voice dialogue between the user and the AI application.
[0176] In one example, after determining the interaction quality, network device 1 can also display the interaction quality. Specifically, the user can control network device 1 to display the interaction quality by entering commands via the command line. In this scenario, network device 1 can also be configured with commands for displaying the interaction quality. This ensures that when the user enters a command to display the interaction quality via the command line, network device 1 can correctly respond to the command and thus display the interaction quality. Displaying the interaction quality refers to displaying information representing that interaction quality.
[0177] In another example, after determining the interaction quality, network device 1 can send indication information to the control and management entity. This indication information indicates the interaction quality, allowing the control and management entity to perform corresponding processing measures based on the indication information. For example, after receiving the indication information, the control and management entity can display the interaction quality. The control and management entity mentioned in this embodiment can be, for example, a device running a network management system (NMS), or a controller. The control and management entity can be a functional module that implements control and / or management functions, or a physical entity running relevant functional modules. For example, the physical entity can be a server with relevant software installed, which implements the functions of the control and management entity. This embodiment does not impose specific limitations.
[0178] In one example, considering that in practical applications, if the audio data sent by the terminal device to the AI application is lost, it may affect the AI application's correct understanding of the semantics of the user's voice data. Consequently, this will affect the AI application's inference efficiency, thereby reducing the efficiency of the AI application in returning corresponding responses to the user, and thus affecting the interaction quality of real-time voice dialogue between the user and the AI application. Therefore, in this application, to ensure the interaction quality of real-time voice dialogue between the user and the AI application, the network device 1 can also execute the method 200 shown in FIG3, which includes S201-S202. FIG3 is a flowchart illustrating a communication method 200 provided in an embodiment of this application.
[0179] S201: Obtain first media data sent by the terminal device regarding real-time voice dialogue between the user and the AI application, wherein the first media data includes first audio data.
[0180] In one example, the first audio data is key audio data from a real-time voice conversation between the user and the AI application. Alternatively, the first audio data is key audio data within the uplink traffic sent by the terminal device to the AI application. For a detailed explanation of key audio data within the uplink traffic, please refer to the preceding description; it will not be repeated here.
[0181] In this scenario, before executing S201, network device 1 can also execute the following steps A1-A2.
[0182] Step A1: Receive second audio data sent by the terminal device, which represents the real-time voice dialogue between the user and the AI application. The second audio data includes non-voice data and the first audio data.
[0183] Step A2: Determine the first audio data in the second audio data.
[0184] As described above, the uplink traffic sent by the terminal device to the AI application includes key voice data and non-voice data. Therefore, network device 1 can receive the second audio data sent by the terminal device. The second audio data includes not only the first audio data but also non-voice data. After receiving the second audio data, network device 1 can determine which data in the second audio data is key voice data. In other words, after receiving the second audio data, network device 1 can determine the first audio data within the second audio data, so as to extract the first audio data from the second audio data. In a specific example, network device 1 can input the traffic characteristics of the second audio data into the AI model to determine the first audio data. Regarding the AI model, traffic characteristics, and how network device 1 determines the first audio data after inputting the traffic characteristics of the second audio data into the AI model, please refer to the relevant descriptions above; they will not be repeated here.
[0185] In yet another example, the first audio data includes not only the key voice data from the real-time voice dialogue between the user and the AI application, but also non-voice data. In this scenario, network device 1 can use the audio data received from the terminal device as the first audio data.
[0186] In another example, considering that when a user and an AI application engage in real-time voice conversations, the data sent by the terminal device to the AI application may include not only audio data but also video data. Therefore, the first media data includes not only the first audio data but also the video data corresponding to the first audio data.
[0187] In a scenario where the first audio data includes non-speech data and key speech data from a real-time voice dialogue between the user and the AI application, the first video data includes invalid video and valid video corresponding to the key speech data from the real-time voice dialogue between the user and the AI application. In this scenario, network device 1 can use video data received from the terminal device as the first video data.
[0188] In real-time voice dialogue scenarios, AI applications typically interpret video data that corresponds to the voice data. Video data unrelated to the voice data is generally left unprocessed. For example, when a user is having a real-time voice conversation with an AI application, their device's camera is on, capturing video to send to the application. If the user points at an object and says, "Please identify what's in the image," the AI application will analyze the video captured during the user's speaking time to identify the objects within that image. Video sent to the AI application during periods when the user is not speaking is not processed. Therefore, valid video corresponding to key voice data refers to video within the same timeframe as the key voice data. For example, if the key voice data is the audio data within the timeframe t1 to t2, then the valid video is the video data within the timeframe t1 to t2. Conversely, invalid video refers to any video sent by the device other than valid video.
[0189] In a scenario where the first audio data is key audio data from a real-time voice dialogue between the user and the AI application, the first video data is valid video corresponding to the key audio data from the real-time voice dialogue between the user and the AI application. In this case, before executing S201, network device 1 may also execute the following steps B1-B3.
[0190] Step B1: Receive second media data sent by the terminal device for real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data of the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video.
[0191] In this application, the second audio data and the second video data sent by the terminal device to the network device 1 are two independent data streams. The second audio data corresponds to an audio stream, and the second video data corresponds to a video stream. Similar to the audio stream, each packet in the video stream also carries a timestamp of its transmission.
[0192] Step B2: Determine the key voice data in the second audio data to obtain the first audio data.
[0193] After receiving the second audio data, network device 1 can determine which data in the second audio data are key voice data. In other words, after receiving the second audio data, network device 1 can identify the first audio data in the second audio data, so as to extract the first audio data from the second audio data. For the specific implementation of step B2, please refer to the previous description of step A2; it will not be repeated here.
[0194] Step B3: Based on the first audio data, determine the first video data in the second video data.
[0195] After determining the first audio data, the video data that is in the same time period as the first audio data can be identified as the first video data based on the timestamps of each message in the second video data. Then, the first video data can be extracted from the second video data.
[0196] S202: When the interaction quality of the real-time voice dialogue does not meet the requirements, the first media data and target data are sent to the AI application, and the target data is used to protect the first media data from packet loss.
[0197] In this application, if the interaction quality of the real-time voice dialogue does not meet the requirements, network device 1 can perform packet loss protection on the first media data, thereby enabling the AI application to receive the first media data as completely as possible, thus improving the efficiency of the AI application in reasoning and, correspondingly, improving the interaction quality of the real-time voice dialogue. The "interaction quality not meeting the requirements" mentioned here can mean that the interaction quality is poor or medium. Alternatively, in a scenario where interaction quality is equivalent to interaction latency, the "interaction quality not meeting the requirements" can mean that the interaction latency is greater than or equal to a certain threshold.
[0198] In this application, network device 1 can use target data to protect the first media data from packet loss.
[0199] In one example, the target data may be the first media data. For example, when the solution of this application embodiment is applied to the scenario shown in Figure 1a or Figure 1b, the target data may be the first media data. In this scenario, after executing S201 and before executing S202, network device 1 can copy the first media data to obtain multiple sets of first media data. For example, network device 1 copies the first media data once, thereby obtaining two sets of first media data. In this scenario, when the receiving device on the AI application side receives multiple sets of first media data, it can discard duplicate packets, thereby achieving packet loss protection for the first media data. Refer to Figure 4a for understanding. Figure 4a is a schematic diagram of an exemplary application scenario provided by an embodiment of this application. Assume that the terminal device sends four packets, namely packet 1, packet 2, packet 3, and packet 4. Network device 1 copies these four packets to obtain two sets of four packets, and sends these two sets of four packets (i.e., eight packets) to the AI application. Assuming that during transmission of these eight messages, if message 2 is lost, the receiving device on the application side can recover it using the other message 2, thus ensuring that the AI application receives the complete first media data. Specifically:
[0200] For the AI-side receiving device, if multiple packets receive audio data and all three parameters (timestamp, sequence number, and packet priority) are identical, the duplicate packets are discarded. Similarly, if multiple packets receive video data and their sequence numbers are identical, the duplicate packets are discarded.
[0201] In another example, the target data may be verification information obtained by FEC encoding the first media data. For example, when the solution of this embodiment is applied to the scenario shown in FIG1b, the target data may be verification information obtained by FEC encoding the first media data. In this scenario, network device 1 can perform FEC encoding on the first media data to obtain the target data after executing S201 and before executing S202. In this case, after receiving the first media data and the target data, network device 2 can use the target data to recover the first media data, thereby achieving packet loss protection for the first media data. Refer to FIG4b for understanding. FIG4b is a schematic diagram of another exemplary application scenario provided by the embodiment of this application. Assume that the terminal device sends four messages, namely message 1, message 2, message 3 and message 4. Network device 1 performs FEC encoding on these four messages to obtain target data R1, R2 and R3, and sends message 1, message 2, message 3, message 4, R1, R2 and R3 to the AI application. Network device 1 can use a generator matrix to calculate R1, R2, and R3 from the aforementioned four packets. Assuming packet 2 is lost, network device 2 can use R1, R2, and R3 to recover the lost packet 2. Specifically, network device 2 can use a recovery matrix to calculate the lost packet 2 from the three received packets and R1, R2, and R3.
[0202] Next, we will introduce the training methods for the AI models mentioned above.
[0203] Referring to Figure 5a, this figure is a flowchart illustrating an AI model training method 300 provided in an embodiment of this application. Method 300 can be executed by a server, which may include, for example, multiple processing units (PUs), including but not limited to, graphics processing units (GPUs) or neural network processing units (NPUs). After the server executes the method shown in Figure 5a to train the AI model, the AI model can be deployed on the aforementioned network device 1.
[0204] Method 300 includes the following S301-S302.
[0205] S301: Obtain training samples, the training samples including: traffic characteristics of the training audio stream and the label of each message in the training audio stream, the label of each message indicating whether each message includes voice data.
[0206] S302: Based on the training samples, the AI model is trained and used to identify whether the message includes voice data.
[0207] In this application, the training audio stream can be a historical audio stream of a user's real-time voice conversation with an AI application, or an audio stream obtained from other scenarios. This application does not impose any specific limitations on the implementation.
[0208] In one example, the training audio stream can include three types of audio streams: the first type includes both speech and non-speech data; the second type includes only speech data; and the third type includes only non-speech data.
[0209] Regarding AI model training, it's important to note that the training process is multi-round iterative, requiring parameter updates to the AI model at each iteration. The process of any given iteration is as follows:
[0210] The AI model obtained from the previous iteration is used to process the traffic features of the input training audio stream to predict the label of each message in the training audio stream. Then, the predicted labels are compared with the actual labels of each message in the training audio stream, and the parameters of the AI model are updated based on the comparison results. As a concrete example, a loss function can be calculated based on the predicted labels and the actual labels of each message in the training audio stream, and the parameters of the AI model can be updated based on the loss function.
[0211] In one example, once the number of iterations of the AI model reaches a certain number, the training of the large AI model is stopped, and the model obtained from the last iteration is used as the training result.
[0212] In another example, training of the large AI model can be stopped when the aforementioned loss function meets certain conditions, and the model obtained when the loss function meets certain conditions can be used as the training result. The condition mentioned here, that the loss function meets certain conditions, means that the accuracy of the AI model's prediction of whether the message includes speech data, meets a certain requirement.
[0213] Currently, the quality of real-time voice interactions between users and AI applications may not meet user needs. For example, after a user asks a question to an AI model via voice, it may take a long time to receive a response from the AI application. One possible reason for this is packet loss in the media data sent by the user to the AI application through their terminal device. This prevents the AI model from correctly understanding the user's intent, thus affecting the inference efficiency of the AI application.
[0214] To address the aforementioned problems, this application provides a communication method, which will now be described in conjunction with the accompanying drawings.
[0215] See Figure 5b, which is a flowchart of a communication method 400 provided in this application.
[0216] Method 400 can be applied to the scenario shown in Figure 1a or Figure 1b, wherein method 400 can be executed by network device 1 shown in Figure 1a or Figure 1b.
[0217] Method 400 includes the following S401-S403.
[0218] S401: Obtain first media data sent by the terminal device for real-time voice dialogue between the user and the AI application, wherein the first media data includes first audio data.
[0219] In one example, the first audio data is key audio data from a real-time voice dialogue between the user and the AI application. This key audio data can be all of the user's audio data or only a portion of it; this embodiment does not impose a specific limitation. When the key audio data is only a portion of the user's audio data, that portion can reflect the complete semantics of the user's total audio data.
[0220] In this scenario, before executing S401, network device 1 can also execute steps A1-A2 in the aforementioned method 200.
[0221] In scenarios where the first audio data is key audio data for a real-time voice dialogue between the user and the AI application, the first video data is a valid video corresponding to the key audio data for the real-time voice dialogue between the user and the AI application. In this case, before executing S401, network device 1 may also execute steps B1-B3 of the aforementioned method 200 to extract the first video data from the second video data.
[0222] As described above, traffic characteristics are used to determine whether a message carries voice data. Therefore, the traffic characteristics can be features that distinguish between messages carrying voice data and messages not carrying voice data. Alternatively, the traffic characteristics can be features that distinguish whether voice data is present. For more information on traffic characteristics, please refer to the relevant descriptions above; they will not be repeated here.
[0223] S402: Copy the first media data to obtain multiple first media data.
[0224] S403: Send the plurality of first media data to the AI application.
[0225] After acquiring the first media data, network device 1 can copy the first media data to obtain multiple copies. For example, network device 1 can copy the first media data once to obtain multiple copies. Further, network device 1 sends the multiple copies of the first media data to the AI application. In this application, since multiple copies of the first media data are sent to the AI application, when some packets in a certain first media data are lost, they can be recovered using corresponding packets from other first media data. This allows the AI application to receive the complete first media data, which in turn helps the AI application correctly understand the user's voice data, thereby improving the inference efficiency of the AI application and ensuring the interaction quality of real-time voice dialogue between the user and the AI application. Refer to Figure 4a for further understanding; for Figure 4a, please refer to the relevant description above, which will not be repeated here.
[0226] In one example, considering that network quality issues might lead to packet loss in the first media data sent by network device 1 to the AI application, consequently affecting the quality of real-time voice interaction between the user and the AI application, network device 1 can duplicate the first media data when network quality is unsatisfactory. This allows the receiving device on the application side to recover the lost packets based on multiple sets of first media data.
[0227] In another example, if the interaction quality of real-time voice dialogue between the user and the AI application does not meet the requirements, it may be due to packet loss in the uplink traffic sent by the user to the AI application. Therefore, when the interaction quality of real-time voice dialogue between the user and the AI application does not meet the requirements, the network device 1 can copy the first media data, so that the receiving device on the application side can recover the lost packets based on multiple first media data, thereby improving the interaction quality of real-time voice dialogue between the user and the AI application.
[0228] In another example, if both network quality and the interaction quality of real-time voice dialogue fail to meet requirements, it indicates that low network quality may be causing packet loss in the uplink traffic sent by the user to the AI application. This further leads to the failure of the interaction quality for real-time voice dialogue between the user and the AI application. Therefore, when both interaction quality and network quality fail to meet requirements, network device 1 can copy the first media data. This allows the receiving device on the application side to recover the lost packets based on multiple sets of first media data, thereby improving the interaction quality of real-time voice dialogue between the user and the AI application.
[0229] The network quality mentioned here can refer to the network quality between network device 1 and the AI application. In one example, network device 1 can test the network quality between itself and the AI application through network testing.
[0230] To address the issue that the interaction quality of real-time voice dialogue between users and AI applications may not meet user needs, this application also provides an alternative communication method. The communication method provided in this application will be described below with reference to the accompanying drawings.
[0231] Referring to Figure 5c, this figure is a flowchart illustrating a communication method 500 provided in an embodiment of this application.
[0232] Method 500 can be executed by the target network device. In one example, the target network device is network device 1 as shown in Figure 1a or Figure 1b. In another example, the target network device is network device 2 as shown in Figure 1b.
[0233] Method 500 includes the following S501-S502.
[0234] S501: Receive first media data for a real-time voice dialogue between the user and the AI application, the first media data including first audio data.
[0235] In one example, if the method shown in Figure 5c is executed by network device 1, that is, the target network device is network device 1, then in a specific implementation, S501 can be receiving the first media data sent by the terminal device.
[0236] In another example, if the method shown in Figure 5c is executed by network device 2, that is, the target network device is network device 2, then in a specific implementation, S501 may be receiving the first media data sent by the AI application.
[0237] In one example, the first audio data is key audio data for real-time voice dialogue between the user and the AI application. Specifically, when the first audio data is audio data sent from the terminal device to network device 1, the key audio data for real-time voice dialogue between the user and the AI application can be all of the user's voice data or a portion of the user's voice data; this embodiment does not impose a specific limitation. When the key audio data is a portion of the user's voice data, this portion of voice data can reflect the complete semantics of the user's entire voice data. Similarly, when the first audio data is audio data sent from the AI application to network device 2, the key audio data for real-time voice dialogue between the user and the AI application can be all of the AI application's voice data or a portion of the AI application's voice data; this embodiment does not impose a specific limitation. When the key audio data is a portion of the AI application's voice data, this portion of voice data can reflect the complete semantics of the AI application's entire voice data.
[0238] In this scenario, before executing S501, the target network device can also execute steps A1-A2 in the aforementioned method 200.
[0239] In scenarios where the first audio data is key audio data from a real-time voice dialogue between the user and the AI application, the first video data is a valid video corresponding to the key audio data from the real-time voice dialogue between the user and the AI application. In this case, before executing S501, the target network device may also execute steps B1-B3 of the aforementioned method 200 to extract the first video data from the second video data.
[0240] As described above, traffic characteristics are used to determine whether a message carries voice data. Therefore, the traffic characteristics can be features that distinguish between messages carrying voice data and messages not carrying voice data. Alternatively, the traffic characteristics can be features that distinguish whether voice data is present. For more information on traffic characteristics, please refer to the relevant descriptions above; they will not be repeated here.
[0241] S502: When the network quality and / or the interaction quality of real-time voice dialogue do not meet the requirements, the first media data and target data are sent to the receiving device, wherein the target data is used to protect the first media data from packet loss.
[0242] In one example, considering that insufficient network quality may lead to packet loss of the first media data, consequently affecting the interaction quality of real-time voice dialogue between the user and the AI application, the target network device can, in one example, protect the first media data from packet loss when network quality is insufficient, thereby improving the interaction quality of real-time voice dialogue between the user and the AI application.
[0243] In another example, considering that if the interaction quality of real-time voice dialogue between the user and the AI application does not meet the requirements, it may be due to packet loss in the uplink traffic sent by the user to the AI application, or packet loss in the downlink traffic sent by the AI application to the terminal device. Therefore, the target network device can perform packet loss protection on the first media data when the interaction quality of real-time voice dialogue between the user and the AI application does not meet the requirements, in order to improve the interaction quality of real-time voice dialogue between the user and the AI application.
[0244] In another example, if both network quality and the interaction quality of real-time voice dialogue fail to meet requirements, it indicates that low network quality may be causing packet loss in the aforementioned uplink or downlink traffic. Furthermore, this results in the interaction quality between the user and the AI application failing to meet requirements. Therefore, the target network device can perform packet loss protection on the first media data when both interaction quality and network quality fail to meet requirements, thereby improving the interaction quality of real-time voice dialogue between the user and the AI application.
[0245] In this application, the target network device can use target data to protect the first media data from packet loss. Specifically, the target network device can send the first media data and the target data to the receiving device. In one example, if the target network device is network device 1, then the receiving device is an AI application. If the target network device is network device 2, then the receiving device is a terminal device.
[0246] In one example, the target data may be the first media data. For instance, when the solution of this embodiment is applied to the scenario shown in Figure 1a or Figure 1b, the target data may be the first media data. In this scenario, after executing S501 and before executing S502, the target network device may copy the first media data to obtain multiple copies of the first media data. For example, the target network device may copy the first media data once to obtain two copies of the first media data. In this scenario, if the receiving device receives multiple copies of the first media data, it may discard duplicate packets, thereby achieving packet loss protection for the first media data. Refer to Figure 4a for further understanding; for Figure 4a, please refer to the relevant description above, which will not be repeated here.
[0247] In another example, the target data may be verification information obtained by FEC encoding the first media data. For example, when the solution of this embodiment is applied to the scenario shown in FIG1b, the target data may be verification information obtained by FEC encoding the first media data. In this scenario, the target network device can perform FEC encoding on the first media data to obtain the target data after executing S501 and before executing S502. In this case, after receiving the first media data and the target data, the receiving network device can use the target data to recover the first media data, thereby achieving packet loss protection for the first media data. Refer to FIG4b for understanding. Regarding FIG4a and FIG4b, please refer to the relevant description above, which will not be repeated here.
[0248] The network quality mentioned in this embodiment can refer to the network quality between network device 1 and the AI application. In one example, network device 1 can detect the network quality between itself and the AI application through network testing. In another example, if the method shown in Figure 5c is executed by network device 2, then after determining the network quality, network device 1 can send the network quality information to network device 2. Specifically, network device 1 can send indication information indicating the network quality to network device 2.
[0249] In one example, the interaction quality of a real-time voice dialogue between a user and an AI application can be determined by network device 1. In another example, if method 500 shown in Figure 5c is executed by network device 2, then after determining the interaction quality, network device 1 can send the interaction quality to network device 2. Specifically, network device 1 can send indication information indicating the interaction quality to network device 2.
[0250] In the method shown in Figure 5b or Figure 5c, network device 1 can use the method shown in Figure 2 to determine the interaction quality of real-time voice dialogue between the user and the AI application.
[0251] The communication method provided by the embodiments of this application has been described above. Next, with reference to Figures 6 and 7, two possible implementation schemes of this application will be introduced.
[0252] Referring to Figure 6, Figure 6 is a schematic diagram of a communication method provided by an embodiment of this application. The method shown in Figure 6 is applied to the application scenarios shown in Figure 1a or Figure 1b.
[0253] As shown in Figure 6, network device 1 includes: a critical voice data recognition module, an interaction quality measurement module, a network quality measurement module, a critical information protection and control module, and a replication module. Among them:
[0254] The key speech data recognition module may include, for example, an AI model, which is used to implement the key speech data recognition function.
[0255] During real-time voice dialogue interaction between the user and the AI application, the network device 1 acquires real-time uplink and downlink traffic, and uses a key voice data recognition module to determine the end time of key voice data in the uplink traffic and the start time of key voice data in the downlink traffic. This end time and start time are then sent to the interaction quality measurement module, which determines the interaction quality of the real-time voice dialogue between the user and the AI application.
[0256] The network quality measurement module measures the network quality between itself and the AI application in real time through network dialing.
[0257] The interaction quality measured by the interaction quality measurement module and the network quality measured by the network quality measurement module are both sent to the critical information protection and control module.
[0258] When the interaction quality and network quality do not meet the requirements, the critical information protection control module will activate the critical information protection function. Specifically, it will activate the copy module to protect the critical information.
[0259] Accordingly, the copying module copies key voice data (or key voice data and valid video) of the user's interaction with the AI application. Correspondingly, the receiving device on the application side recovers lost packets during transmission and removes duplicate received packets.
[0260] Regarding the process shown in Figure 6, an example is given below:
[0261] The user begins real-time voice interaction with the AI application. For the first three real-time voice interactions, network device 1 determines that the interaction quality and network quality are both unsatisfactory. Therefore, starting from the fourth real-time voice interaction, network device 1 copies the key voice data (or key voice data and valid video) sent by the terminal device, thereby protecting the key voice data from packet loss and ensuring that the AI application receives complete key voice data. From the fourth real-time voice interaction onwards, the interaction quality between the user and the AI application improves.
[0262] Referring to Figure 7, which is a schematic diagram of a communication method provided in an embodiment of this application, the method shown in Figure 7 is applied to the application scenario shown in Figure 1b.
[0263] As shown in Figure 7, network device 1 includes: a critical voice data recognition module, an interaction quality measurement module, a network quality measurement module, a critical information protection and control module, and an FEC module. Among them:
[0264] The key speech data recognition module may include, for example, an AI model, which is used to implement the key speech data recognition function.
[0265] During real-time voice dialogue interaction between the user and the AI application, the network device 1 acquires real-time uplink and downlink traffic, and uses a key voice data recognition module to determine the end time of key voice data in the uplink traffic and the start time of key voice data in the downlink traffic. This end time and start time are then sent to the interaction quality measurement module, which determines the interaction quality of the real-time voice dialogue between the user and the AI application.
[0266] The network quality measurement module measures the network quality between itself and the AI application in real time through network dialing.
[0267] The interaction quality measured by the interaction quality measurement module and the network quality measured by the network quality measurement module are both sent to the critical information protection and control module.
[0268] When the interaction quality and network quality do not meet the requirements, the critical information protection control module will activate the critical information protection function. Specifically, the FEC module will be activated to protect the critical information.
[0269] Accordingly, the FEC module performs FEC encoding on the key voice data (or key voice data and valid video) of the user's interaction with the AI application to obtain verification information. Correspondingly, network device 2 uses the verification information to recover lost packets during transmission.
[0270] Regarding the process shown in Figure 7, an example is given below:
[0271] The user begins real-time voice interaction with the AI application. For the first three real-time voice interactions, network device 1 determines that the interaction quality and network quality are both unsatisfactory. Therefore, starting from the fourth real-time voice interaction, network device 1 performs FEC encoding on the key voice data (or key voice data and valid video) sent by the terminal device, thereby protecting the key voice data from packet loss and ensuring that the AI application receives complete key voice data. From the fourth real-time voice interaction onwards, the interaction quality between the user and the AI application improves.
[0272] Based on the methods provided in the above embodiments, this application also provides a corresponding apparatus. Next, with reference to the accompanying drawings, the apparatus provided in this application will be described.
[0273] Referring to Figure 8a, this figure is a schematic diagram of the structure of an information processing device provided in an embodiment of this application.
[0274] In one example, the information processing device 800 shown in FIG8a is applied to network device 1 to execute the information processing method performed by network device 1 provided in the above embodiments, for example, to execute the method performed by network device 1 as shown in FIG2 or FIG3.
[0275] As shown in Figure 8a, the device 800 includes a processing unit 801 and a transceiver unit 802, wherein the transceiver unit 802 is optional. The processing unit 801 is used to perform data processing operations, and the transceiver unit 802 is used to perform transceiver operations, wherein the transceiver operations include receiving operations and / or sending operations. In one example, the transceiver unit 802 includes a receiving unit and / or a sending unit, wherein the receiving unit is used to perform receiving operations, and the sending unit is used to perform sending operations.
[0276] The processing unit 801 is configured to: input the traffic characteristics of uplink traffic and downlink traffic into an artificial intelligence (AI) model to determine the interaction latency of real-time voice dialogue between the user and the AI application, wherein the uplink traffic includes the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes the audio stream sent by the AI application to the terminal device; and determine the interaction quality of real-time voice dialogue between the user and the AI application based on the interaction latency.
[0277] In one possible implementation, the step of inputting the uplink traffic characteristics and downlink traffic characteristics into an artificial intelligence (AI) model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: inputting the uplink traffic characteristics into the AI model to determine key voice data in the uplink traffic; inputting the downlink traffic characteristics into the AI model to determine key voice data in the downlink traffic; and determining the interaction latency based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0278] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0279] In one possible implementation, the user and the AI application engage in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the interaction latency is the average latency of the multiple real-time voice conversations.
[0280] In one possible implementation, determining the interaction quality of the real-time voice dialogue between the user and the AI application based on the interaction latency includes: comparing the interaction latency with a latency threshold to determine the interaction quality.
[0281] In one possible implementation, the processing unit 801 is further configured to: configure a command for displaying the interaction quality.
[0282] In one possible implementation, the sending unit is configured to send indication information to the control and management entity, the indication information being used to indicate the interaction quality.
[0283] In one possible implementation, the apparatus further includes: an acquisition unit, configured to acquire first media data sent by the terminal device for real-time voice dialogue between the user and the AI application, the first media data including first audio data; and a sending unit, further configured to send the first media data and target data to the AI application when the interaction quality of the real-time voice dialogue does not meet the requirements, the target data being used for packet loss protection of the first media data.
[0284] In one example, the acquisition unit is a receiving unit. In another example, the acquisition unit is the processing unit 801, in which case the processing unit 801 is also used to acquire the first media data.
[0285] In one possible implementation, the processing unit 801 is further configured to: perform forward error correction (FEC) encoding on the first media data to obtain the target data before sending the first media data and the target data to the AI application; or, copy the first media data to obtain the target data before sending the first media data and the target data to the AI application.
[0286] In one possible implementation, the first audio data is key audio data from a real-time voice conversation between the user and the AI application.
[0287] In one possible implementation, the receiving unit is configured to receive second audio data sent by the terminal device, representing a real-time voice dialogue between the user and the AI application, wherein the second audio data includes non-voice data and the first audio data; the processing unit 801 is further configured to determine the first audio data in the second audio data.
[0288] In one possible implementation, determining the first audio data in the second audio data includes: inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0289] In one possible implementation, the first media data further includes first video data corresponding to the first audio data.
[0290] In one possible implementation, the first video data is a valid video corresponding to key voice data of a real-time voice conversation between the user and the AI application.
[0291] In one possible implementation, the receiving unit is configured to receive second media data sent by the terminal device, representing a real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data representing the voice dialogue between the user and the AI application. The second video data includes invalid video and valid video. The processing unit 801 is further configured to determine the key voice data in the second audio data to obtain the first audio data; and to determine the first video data in the second video data based on the first audio data.
[0292] In another example, the information processing apparatus 800 shown in FIG8a is applied to network device 1 to execute the communication method performed by network device 1 as provided in the above embodiments, for example, executing the method performed by network device 1 as shown in FIG5b. In this case:
[0293] The processing unit 801 is used to acquire first media data sent by the terminal device for real-time voice dialogue between the user and the artificial intelligence (AI) application, the first media data including first audio data; and to copy the first media data to obtain multiple first media data.
[0294] The sending unit is used to send the plurality of first media data to the AI application.
[0295] In one possible implementation, the first audio data is key audio data from a real-time voice conversation between the user and the AI application.
[0296] In one possible implementation, the receiving unit is configured to receive second audio data sent by the terminal device, the second audio data including non-voice data and the first audio data; the processing unit 801 is further configured to determine the first audio data in the second audio data.
[0297] In one possible implementation, determining the first audio data in the second audio data includes: inputting the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0298] One possible implementation includes one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0299] In one possible implementation, the first media data further includes first video data corresponding to the first audio data.
[0300] In one possible implementation, the first video data is a valid video corresponding to key voice data of a real-time voice conversation between the user and the AI application.
[0301] In one possible implementation, the receiving unit is configured to receive second media data sent by the terminal device. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data of the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. The processing unit 801 is further configured to determine the key voice data in the second audio data to obtain the first audio data; and to determine the first video data in the second video data based on the first audio data.
[0302] In one possible implementation, copying the first media data includes: copying the first media data when the network quality and / or the interaction quality of the real-time voice dialogue do not meet the requirements.
[0303] In one possible implementation, the processing unit 801 is further configured to: input the traffic characteristics of uplink traffic and the traffic characteristics of downlink traffic into the AI model respectively, to determine the interaction latency of real-time voice dialogue between the user and the AI application, wherein the uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device; and determine the interaction quality based on the interaction latency.
[0304] In one possible implementation, the step of inputting the uplink traffic characteristics and downlink traffic characteristics into an AI speech recognition model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: inputting the uplink traffic characteristics into the AI model to determine key voice data in the uplink traffic; inputting the downlink traffic characteristics into the AI model to determine key voice data in the downlink traffic; and determining the interaction latency based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0305] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0306] In one possible implementation, the user and the AI application engage in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the interaction latency is the average latency of the multiple real-time voice conversations.
[0307] In one possible implementation, determining the interaction quality based on the interaction latency includes comparing the interaction latency with a latency threshold to determine the interaction quality.
[0308] Referring to Figure 8b, this figure is a schematic diagram of the structure of a communication device provided in an embodiment of this application. The communication device 800 shown in Figure 8b is applied to a target network device and is used to execute the method provided by the target network device in the above embodiments, for example, to execute the method shown in Figure 5c.
[0309] As shown in Figure 8b, the device 800 includes a receiving unit 803 and a transmitting unit 804.
[0310] The receiving unit 803 is used to receive first media data for real-time voice dialogue between the user and the artificial intelligence (AI) application, the first media data including first audio data.
[0311] The sending unit 804 is used to send the first media data and target data to the receiving device when the network quality and / or the interaction quality of real-time voice dialogue do not meet the requirements. The target data is used to protect the first media data from packet loss.
[0312] In one possible implementation, the apparatus further includes a processing unit configured to: perform forward error correction (FEC) encoding on the first media data to obtain the target data before sending the first media data and the target data to the receiving device; or, copy the first media data to obtain the target data before sending the first media data and the target data to the receiving device.
[0313] In one possible implementation, the receiving unit 803 is configured to: receive the first media data sent by the terminal device; and the sending unit 804 is configured to: send the first media data and the target data to the AI application.
[0314] In one possible implementation, the receiving unit 803 is configured to: receive the first media data sent by the AI application; and the sending unit 804 is configured to: send the first media data and the target data to the terminal device.
[0315] In one possible implementation, the processing unit of the device is further configured to: input the traffic characteristics of uplink traffic and the traffic characteristics of downlink traffic into the AI model respectively, to determine the interaction latency of real-time voice dialogue between the user and the AI application, wherein the uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device; and determine the interaction quality based on the interaction latency.
[0316] In one possible implementation, the step of inputting the uplink traffic characteristics and downlink traffic characteristics into an AI speech recognition model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: inputting the uplink traffic characteristics into the AI model to determine key voice data in the uplink traffic; inputting the downlink traffic characteristics into the AI model to determine key voice data in the downlink traffic; and determining the interaction latency based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
[0317] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0318] In one possible implementation, the user and the AI application engage in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the interaction latency is the average latency of the multiple real-time voice conversations.
[0319] In one possible implementation, determining the interaction quality based on the interaction latency includes comparing the interaction latency with a latency threshold to determine the interaction quality.
[0320] In one possible implementation, the first audio data is key audio data from a real-time voice conversation between the user and the AI application.
[0321] In one possible implementation, the receiving unit 803 is further configured to receive second audio data of a real-time voice dialogue between the user and the AI application, the second audio data including non-voice data and the first audio data; the processing unit is further configured to determine the first audio data in the second audio data.
[0322] In one possible implementation, the processing unit is configured to: input the traffic characteristics of the second audio data into an AI model to determine the first audio data.
[0323] In one possible implementation, the first media data further includes first video data corresponding to the first audio data.
[0324] In one possible implementation, the first video data is a valid video corresponding to key voice data of a real-time voice conversation between the user and the AI application.
[0325] In one possible implementation, the receiving unit 803 is further configured to receive second media data for real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data for the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. The processing unit is further configured to determine the key voice data in the second audio data to obtain the first audio data. Based on the first audio data, the processing unit determines the first video data in the second video data.
[0326] Referring to Figure 9, this figure is a schematic diagram of the structure of an AI model training device provided in an embodiment of this application. The AI model training device 900 shown in Figure 9 is used to execute the AI model training method provided in the above embodiments, for example, to execute the AI model training method shown in Figure 5a.
[0327] The device 900 shown in Figure 9 includes an acquisition unit 901 and a processing unit 902.
[0328] The acquisition unit 901 is used to acquire training samples, which include: traffic characteristics of the training audio stream and a label for each message in the training audio stream, wherein the label for each message indicates whether the message includes voice data.
[0329] The processing unit 902 is used to train an AI model based on the training samples, and the AI model is used to identify whether the message includes voice data.
[0330] In one possible implementation, the traffic characteristics include one or more of the following: message size, packet sending rate, and traffic burst characteristics.
[0331] For the specific implementation of each unit of the devices 800 and 900, please refer to the description of the relevant methods provided in the embodiments of this application above, which will not be repeated here.
[0332] Furthermore, this application embodiment also provides a communication device 1000, as shown in FIG10, which is a structural schematic diagram of a communication device provided in this application embodiment. The communication device 1000 includes a communication interface 1001 and a processor 1002 connected to the communication interface 1001. The communication interface 1001 is optional. The communication device 1000 can be used to execute the methods in the above embodiments, for example, to execute the methods shown in FIG2, FIG3, FIG5a, FIG5b, or FIG5c.
[0333] When the communication device 1000 is used to execute the method shown in FIG2, the communication interface 1001 is used to execute the receiving and / or sending operations in the method shown in FIG2. The processor 1002 is used to execute other operations in the method shown in FIG2 besides the receiving and / or sending operations. For example, the processor 1002 is used to input the traffic characteristics of uplink traffic and downlink traffic into the AI model respectively to determine the interaction latency of real-time voice dialogue between the user and the AI application. The uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device. Based on the interaction latency, the interaction quality of real-time voice dialogue between the user and the AI application is determined.
[0334] Optionally, the communication interface 1001 is used to send indication information to the control and management entity, the indication information being used to indicate the quality of the interaction.
[0335] When the communication device 1000 executes the method shown in FIG3, the communication interface 1001 is used to execute the receiving and / or sending operations in the method shown in FIG3. The processor 1002 is used to execute other operations in the method shown in FIG3 besides the receiving and / or sending operations. For example, the processor 1002 is used to acquire first media data sent by the terminal device for real-time voice dialogue between the user and the AI application, the first media data including first audio data. When the interaction quality of the real-time voice dialogue does not meet the requirements, the communication interface 1001 is used to send the first media data and target data to the AI application, the target data being used for packet loss protection of the first media data.
[0336] When the communication device 1000 is used to execute the method shown in FIG. 5a, the communication interface 1001 is used to execute the receiving and / or sending operations in the method shown in FIG. 5a. The processor 1002 is used to execute other operations in the method shown in FIG. 5a besides the receiving and / or sending operations. For example, the processor 1002 is used to acquire training samples, the training samples including: traffic characteristics of the training audio stream and a tag for each message in the training audio stream, the tag of each message indicating whether each message includes voice data; based on the training samples, the AI model is trained, the AI model being used to identify whether a message includes voice data.
[0337] When the communication device 1000 is used to execute the method shown in FIG. 5b, the communication interface 1001 is used to execute the receiving and / or sending operations in the method shown in FIG. 5b. The processor 1002 is used to execute other operations in the method shown in FIG. 5b besides the receiving and / or sending operations. For example, the processor 1002 is used to acquire first media data sent by the terminal device for a real-time voice dialogue between the user and an artificial intelligence (AI) application, the first media data including first audio data; and to copy the first media data to obtain a plurality of first media data. The communication interface 1001 is used to send the plurality of first media data to the AI application.
[0338] When the communication device 1000 is used to execute the method shown in FIG. 5c, the communication interface 1001 is used to execute the receiving and / or sending operations in the method shown in FIG. 5c. The processor 1002 is used to execute other operations in the method shown in FIG. 5c besides the receiving and / or sending operations. For example: the communication interface 1001 is used to receive first media data for real-time voice dialogue between the user and the artificial intelligence (AI) application, the first media data including first audio data; when the network quality and / or the interaction quality of the real-time voice dialogue does not meet the requirements, the first media data and target data are sent to the receiving device, the target data being used to protect the first media data from packet loss. Optionally, the processor 1002 is used to determine that the network quality and / or the interaction quality of the real-time voice dialogue does not meet the requirements.
[0339] Furthermore, this application embodiment also provides a communication device 1100, as shown in FIG11, which is a schematic diagram of the structure of a communication device provided in this application embodiment. This communication device 1100 can be used to execute the methods in the above embodiments, for example, to execute the methods shown in FIG2, FIG3, FIG5a, FIG5b, or FIG5c.
[0340] As shown in Figure 11, the communication device 1100 may include a processor 1110, a communication interface 1120, and a memory 1130 coupled to the processor 1110. The communication interface 1120 and the memory 1130 are optional.
[0341] The processor mentioned in this application can be one or more processors. When there are multiple processors, the types of processors can be the same or different. A processor can be, for example, a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. A processor can also be one or more processing circuits. A processor can also be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0342] The memory 1130 mentioned in this application may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 1130 may also include combinations of the above types of memory. The memory 1130 may refer to a single memory or may include multiple memories. In one embodiment, the memory 1130 stores computer-readable instructions, which include multiple software modules, such as a sending module 1131, a processing module 1132, and a receiving module 1133. After executing each software module, the processor 1110 can perform corresponding operations according to the instructions of each software module. In this embodiment, the operation performed by a software module actually refers to the operation performed by the processor 1110 according to the instructions of the software module.
[0343] When the communication device 1100 is used to execute the method shown in FIG2, the communication interface 1120 is used to execute the receiving and / or sending operations in the method shown in FIG2. The processor 1110 is used to execute other operations in the method shown in FIG2 besides the receiving and / or sending operations. For example, the processor 1110 is used to input the traffic characteristics of uplink traffic and downlink traffic into the AI model respectively to determine the interaction latency of real-time voice dialogue between the user and the AI application. The uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device. Based on the interaction latency, the interaction quality of real-time voice dialogue between the user and the AI application is determined.
[0344] Optionally, the communication interface 1120 is used to send indication information to the control and management entity, the indication information being used to indicate the quality of the interaction.
[0345] When the communication device 1100 executes the method shown in FIG3, the communication interface 1120 is used to execute the receiving and / or sending operations in the method shown in FIG3. The processor 1110 is used to execute other operations in the method shown in FIG3 besides the receiving and / or sending operations. For example, the processor 1110 is used to acquire first media data sent by the terminal device for real-time voice dialogue between the user and the AI application, the first media data including first audio data. The communication interface 1120 is used to send the first media data and target data to the AI application when the interaction quality of the real-time voice dialogue does not meet the requirements, the target data being used for packet loss protection of the first media data.
[0346] When the communication device 1100 is used to execute the method shown in FIG. 5a, the communication interface 1120 is used to execute the receiving and / or sending operations in the method shown in FIG. 5a. The processor 1110 is used to execute other operations in the method shown in FIG. 5a besides the receiving and / or sending operations. For example, the processor 1110 is used to acquire training samples, the training samples including: traffic characteristics of the training audio stream and a tag for each message in the training audio stream, the tag of each message indicating whether each message includes voice data; based on the training samples, the AI model is trained, the AI model being used to identify whether a message includes voice data.
[0347] When the communication device 1100 is used to execute the method shown in FIG. 5b, the communication interface 1120 is used to execute the receiving and / or sending operations in the method shown in FIG. 5b. The processor 1110 is used to execute other operations in the method shown in FIG. 5b besides the receiving and / or sending operations. For example, the processor 1110 is used to acquire first media data sent by the terminal device for a real-time voice dialogue between the user and an artificial intelligence (AI) application, the first media data including first audio data; and to copy the first media data to obtain a plurality of first media data. The communication interface 1120 is used to send the plurality of first media data to the AI application.
[0348] When the communication device 1100 is used to execute the method shown in FIG. 5c, the communication interface 1120 is used to execute the receiving and / or sending operations in the method shown in FIG. 5c. The processor 1110 is used to execute other operations in the method shown in FIG. 5c besides the receiving and / or sending operations. For example, the communication interface 1120 is used to receive first media data for real-time voice dialogue between the user and the artificial intelligence (AI) application, the first media data including first audio data; when the network quality and / or the interaction quality of the real-time voice dialogue does not meet the requirements, the processor 1110 sends the first media data and target data to the receiving device, the target data being used to protect the first media data from packet loss. Optionally, the processor 1110 is used to determine that the network quality and / or the interaction quality of the real-time voice dialogue does not meet the requirements.
[0349] This application also provides a computer-readable storage medium storing instructions or computer programs that, when executed on a processor, can implement any one or more operations of the methods described in the foregoing embodiments (e.g., the methods shown in Figure 2, Figure 3, Figure 5a, Figure 5b, or Figure 5c).
[0350] This application also provides a computer program product, including a computer program that, when run on a processor, can implement any one or more of the methods described in the foregoing embodiments (e.g., the methods shown in FIG2, FIG3, FIG5a, FIG5b, or FIG5c).
[0351] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0352] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0353] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical business division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0354] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0355] Furthermore, the various business units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software business unit.
[0356] If the integrated unit is implemented as a software business unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0357] Those skilled in the art will recognize that, in one or more of the examples above, the services described in this application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these services can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of computer programs from one place to another. Storage media can be any available medium accessible to general-purpose or special-purpose computers.
[0358] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application.
[0359] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. An information processing method, characterized in that, The method includes: The uplink traffic characteristics and downlink traffic characteristics are respectively input into the artificial intelligence (AI) model to determine the interaction latency of real-time voice dialogue between the user and the AI application. The uplink traffic includes the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes the audio stream sent by the AI application to the terminal device. The interaction quality between the user and the AI application in real-time voice dialogue is determined based on the interaction latency.
2. The method according to claim 1, characterized in that, The step of inputting the uplink traffic characteristics and downlink traffic characteristics into the artificial intelligence (AI) model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: The uplink traffic characteristics are input into the AI model to determine the key voice data in the uplink traffic; The traffic characteristics of the downlink traffic are input into the AI model to determine the key voice data in the downlink traffic; The interaction delay is determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
3. The method according to claim 1 or 2, characterized in that, The traffic characteristics include one or more of the following: Message size, packet sending rate, and traffic burst characteristics.
4. The method according to any one of claims 1-3, characterized in that, The user and the AI application engaged in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the average latency of the multiple real-time voice conversations.
5. The method according to any one of claims 1-4, characterized in that, Determining the interaction quality of the real-time voice dialogue between the user and the AI application based on the interaction latency includes: The interaction latency and latency threshold are compared to determine the interaction quality.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Configure commands for displaying the quality of the interaction.
7. The method according to any one of claims 1-5, characterized in that, The method further includes: Send instruction information to the control and management entity, the instruction information being used to indicate the quality of the interaction.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: The first media data sent by the terminal device for real-time voice dialogue between the user and the AI application is obtained, wherein the first media data includes first audio data. When the interaction quality of real-time voice dialogue does not meet the requirements, the first media data and target data are sent to the AI application, and the target data is used to protect the first media data from packet loss.
9. The method according to claim 8, characterized in that, Before sending the first media data and target data to the AI application, the method further includes: The target data is obtained by performing forward error correction (FEC) encoding on the first media data; or, Copy the first media data to obtain the target data.
10. The method according to claim 8 or 9, characterized in that, The first audio data is key audio data from the real-time voice dialogue between the user and the AI application.
11. The method according to claim 10, characterized in that, The method further includes: The terminal device receives second audio data of a real-time voice dialogue between the user and the AI application, the second audio data including non-voice data and the first audio data; The first audio data is determined from the second audio data.
12. The method according to claim 11, characterized in that, Determining the first audio data in the second audio data includes: The traffic characteristics of the second audio data are input into the AI model to determine the first audio data.
13. The method according to claim 8 or 9, characterized in that, The first media data also includes first video data corresponding to the first audio data.
14. The method according to claim 13, characterized in that, The first video data is a valid video corresponding to the key voice data of the real-time voice dialogue between the user and the AI application.
15. The method according to claim 14, characterized in that, The method further includes: The terminal device receives second media data sent by the terminal device, which is used for real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data of the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. The key voice data in the second audio data is determined to obtain the first audio data; Based on the first audio data, the first video data in the second video data is determined.
16. A method for training an artificial intelligence (AI) model, characterized in that, The method includes: Obtain training samples, which include: traffic features of the training audio stream and a label for each message in the training audio stream, wherein the label for each message indicates whether the message contains voice data; Based on the training samples, an AI model is trained, which is used to identify whether a message includes voice data.
17. The method according to claim 16, characterized in that, The traffic characteristics include one or more of the following: Message size, packet sending rate, and traffic burst characteristics.
18. A method of communication, comprising: The method includes: Acquire first media data sent by the terminal device for a real-time voice dialogue between the user and an artificial intelligence (AI) application, wherein the first media data includes first audio data; Copy the first media data to obtain multiple first media data sets; Send the plurality of first media data to the AI application.
19. The method of claim 18, wherein, The first audio data is key audio data from the real-time voice dialogue between the user and the AI application.
20. The method of claim 19, wherein, The method further includes: Receive second audio data sent by the terminal device, the second audio data including non-voice data and the first audio data; The first audio data is determined from the second audio data.
21. The method according to claim 20, characterized in that, Determining the first audio data in the second audio data includes: The traffic characteristics of the second audio data are input into the AI model to determine the first audio data.
22. The method according to claim 21, characterized in that, The traffic characteristics include one or more of the following: Message size, packet sending rate, and traffic burst characteristics.
23. The method according to claim 18 or 19, characterized in that, The first media data also includes first video data corresponding to the first audio data.
24. The method according to claim 23, characterized in that, The first video data is a valid video corresponding to the key voice data of the real-time voice dialogue between the user and the AI application.
25. The method according to claim 24, characterized in that, The method further includes: The terminal device sends second media data, which includes second audio data and second video data. The second audio data includes non-voice data and key voice data of the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. The key voice data in the second audio data is determined to obtain the first audio data; Based on the first audio data, the first video data in the second video data is determined.
26. The method according to any one of claims 18-25, characterized in that, The copying of the first media data includes: If the network quality and / or the interaction quality of the real-time voice dialogue do not meet the requirements, the first media data is copied.
27. The method according to claim 26, characterized in that, The method further includes: The uplink traffic characteristics and downlink traffic characteristics are respectively input into the AI model to determine the interaction latency of real-time voice dialogue between the user and the AI application. The uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device. The interaction quality is determined based on the interaction latency.
28. The method according to claim 27, characterized in that, The step of inputting the uplink traffic characteristics and downlink traffic characteristics into the AI speech recognition model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: The uplink traffic characteristics are input into the AI model to determine the key voice data in the uplink traffic; The traffic characteristics of the downlink traffic are input into the AI model to determine the key voice data in the downlink traffic; The interaction delay is determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
29. The method according to claim 27 or 28, characterized in that, The traffic characteristics include one or more of the following: Message size, packet sending rate, and traffic burst characteristics.
30. The method according to any one of claims 27-29, characterized in that, The user and the AI application engaged in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the average latency of the multiple real-time voice conversations.
31. The method according to any one of claims 27-30, characterized in that, Determining the interaction quality based on the interaction latency includes: The interaction latency and latency threshold are compared to determine the interaction quality.
32. A communication method, characterized in that, The method includes: Receive first media data of a real-time voice conversation between a user and an artificial intelligence (AI) application, wherein the first media data includes first audio data; When the network quality and / or the interaction quality of real-time voice dialogue do not meet the requirements, the first media data and target data are sent to the receiving device, and the target data is used to protect the first media data from packet loss.
33. The method according to claim 32, characterized in that, Before sending the first media data and target data to the receiving device, the method further includes: The target data is obtained by performing forward error correction (FEC) encoding on the first media data; or, Copy the first media data to obtain the target data.
34. The method according to claim 32 or 33, characterized in that, The first media data for receiving real-time voice dialogue between the user and the AI application includes: receiving the first media data sent by the terminal device; Sending the first media data and the target data to the receiving device includes sending the first media data and the target data to the AI application.
35. The method according to claim 32 or 33, characterized in that, The first media data for receiving real-time voice dialogue between the user and the AI application includes: receiving the first media data sent by the AI application; Sending the first media data and the target data to the receiving device includes sending the first media data and the target data to the terminal device.
36. The method according to any one of claims 32-35, characterized in that, The method further includes: The uplink traffic characteristics and downlink traffic characteristics are respectively input into the AI model to determine the interaction latency of real-time voice dialogue between the user and the AI application. The uplink traffic includes: the audio stream sent by the user to the AI application through the terminal device, and the downlink traffic includes: the audio stream sent by the AI application to the terminal device. The interaction quality is determined based on the interaction latency.
37. The method according to claim 36, characterized in that, The step of inputting the uplink traffic characteristics and downlink traffic characteristics into the AI speech recognition model to determine the interaction latency for real-time voice dialogue between the user and the AI application includes: The uplink traffic characteristics are input into the AI model to determine the key voice data in the uplink traffic; The traffic characteristics of the downlink traffic are input into the AI model to determine the key voice data in the downlink traffic; The interaction delay is determined based on the end time of the key voice data in the uplink traffic and the start time of the key voice data in the downlink traffic.
38. The method according to claim 36 or 37, characterized in that, The traffic characteristics include one or more of the following: Message size, packet sending rate, and traffic burst characteristics.
39. The method according to any one of claims 36-38, characterized in that, The user and the AI application engaged in multiple real-time voice conversations within a certain time period. The interaction latency is either the latency of any single real-time voice conversation between the user and the AI application, or the average latency of the multiple real-time voice conversations.
40. The method according to any one of claims 36-39, characterized in that, Determining the interaction quality based on the interaction latency includes: The interaction latency and latency threshold are compared to determine the interaction quality.
41. The method according to any one of claims 32-40, characterized in that, The first audio data is key audio data from the real-time voice dialogue between the user and the AI application.
42. The method according to claim 41, characterized in that, The method further includes: Receive second audio data of a real-time voice conversation between the user and the AI application, wherein the second audio data includes non-voice data and the first audio data; The first audio data is determined from the second audio data.
43. The method according to claim 42, characterized in that, Determining the first audio data in the second audio data includes: The traffic characteristics of the second audio data are input into the AI model to determine the first audio data.
44. The method according to any one of claims 32-40, characterized in that, The first media data also includes first video data corresponding to the first audio data.
45. The method according to claim 44, characterized in that, The first video data is a valid video corresponding to the key voice data of the real-time voice dialogue between the user and the AI application.
46. The method according to claim 45, characterized in that, The method further includes: Receive second media data of real-time voice dialogue between the user and the AI application. The second media data includes second audio data and second video data. The second audio data includes non-voice data and key voice data of the voice dialogue between the user and the AI application. The second video data includes invalid video and the valid video. The key voice data in the second audio data is determined to obtain the first audio data; Based on the first audio data, the first video data in the second video data is determined.
47. A communication device, characterized in that, The communication device includes multiple units that interact with each other to implement the method according to any one of claims 1-46.
48. A communication device, characterized in that, include: Processor and memory; The memory is used to store instructions; The processor is configured to execute the instructions, causing the communication device to perform the method according to any one of claims 1-46.
49. A computer-readable storage medium, characterized in that, Includes instructions or computer programs that, when executed on a processor, implement the method described in any one of claims 1-46.
50. A computer program product, characterized in that, The computer program product includes instructions or a computer program that, when run on a computer, causes the computer to perform the method described in any one of claims 1-46.