Communication method, system and device, electronic equipment and computer readable storage medium
By predicting future QoS vectors and using a dual-time-window model to pre-schedule audio and video stream processing, the problem of communication lag was solved, resulting in a smooth and accurate user experience that can adapt to complex network environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies, when dealing with communication disruptions and audio/video desynchronization caused by factors such as network fluctuations, suffer from response delays and inaccurate adjustments in post-incident remedial measures. They also fail to anticipate network fluctuation trends in advance, thus impacting user experience.
By predicting future QoS vectors at the sending end, and fusing short-term and long-term communication features based on a dual-time-window prediction model, audio and video stream processing is scheduled in advance to generate target data and send it. The receiving end uses the generation model to recover the communication content.
It enables advance scheduling before network fluctuations occur, avoiding communication disruptions, ensuring a smooth and accurate user experience, and adapting to dynamic communication needs in multiple scenarios.
Smart Images

Figure CN121664786A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, specifically to a communication method, system, device, electronic device, and computer-readable storage medium. Background Technology
[0002] With the widespread adoption of remote work, online education, video conferencing, and social entertainment, real-time communication has become an indispensable basic capability in daily work and life, and its smoothness directly affects communication efficiency and user experience. In actual use, network signal fluctuations and insufficient bandwidth can easily lead to communication lag and audio-video desynchronization. Related technologies adjust by compensating for missing frames after detecting communication lag. However, this reactive approach often takes several seconds from the occurrence of lag to the parameter adjustment taking effect; during this time, lag and screen glitches have already caused irreversible damage to the user experience. Summary of the Invention
[0003] This application provides a communication method, system, device, electronic device, and computer-readable storage medium that can predict and schedule communication quality in advance, thereby ensuring the user's communication experience.
[0004] In a first aspect, embodiments of this application provide a communication method applied at a sending end; the method includes: Predict future QoS vectors based on current communication characteristics; Based on the future QoS vector, the collected audio and / or video streams are processed to obtain the target data; The target data is sent to the receiving end so that the receiving end can input the target data into the generation model to obtain the communication content.
[0005] Secondly, embodiments of this application provide a communication method applied at a receiving end; the method includes: Receive the target data sent by the sending end; The target data is input into the generation model to obtain the communication content; The target data is obtained by the sending end predicting the future QoS vector based on the current communication characteristics, and processing the collected audio stream and / or video stream according to the future QoS vector.
[0006] Thirdly, embodiments of this application provide a communication system, including: The sending end is used to predict the future QoS vector based on the current communication characteristics; process the collected audio stream and / or video stream according to the future QoS vector to obtain target data, and send the target data to the receiving end; The receiving end is used to input the target data into the generation model to obtain the communication content.
[0007] Fourthly, embodiments of this application provide a communication device applied at a transmitting end; the device includes: The prediction module is used to predict future QoS vectors based on current communication characteristics; The processing module is used to process the acquired audio and / or video streams according to the future QoS vector to obtain the target data; The sending module is used to send the target data to the receiving end, so that the receiving end can input the target data into the generation model to obtain the communication content.
[0008] Fifthly, embodiments of this application provide a communication device applied at a receiving end; the device includes: The receiving module is used to receive the target data sent by the sending end; The input module is used to input the target data into the generation model to obtain the communication content; The target data is obtained by the sending end predicting the future QoS vector based on the current communication characteristics, and processing the collected audio stream and / or video stream according to the future QoS vector.
[0009] Sixthly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the communication method described above.
[0010] In a seventh aspect, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the communication method described above.
[0011] Eighthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.
[0012] The embodiments of this application have the following beneficial effects: The sending end can predict the future Quality of Service (QoS) vector in advance based on the current communication characteristics, and schedule the audio and video streams in advance according to the predicted future QoS vector to obtain the target data. When transmitting the target data based on the future network, problems such as lag can be avoided. The receiving end can input the target data into the generation model to obtain the communication content, thereby ensuring that the receiving end can accurately obtain the communication content and achieve effective communication. In this way, by predicting the communication quality in advance and scheduling it in advance, the user's communication experience can be guaranteed. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the steps of a communication method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a dual-time-window prediction model provided in an embodiment of this application; Figure 3 This is a schematic diagram of another step of the communication method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the architecture of a multimodal generation model provided in an embodiment of this application; Figure 5 This is a schematic diagram of the digital watermarking processing flow provided in an embodiment of this application; Figure 6 This is a timing diagram of a communication method provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a communication system provided in an embodiment of this application; Figure 8 This is another schematic diagram of the communication system provided in one embodiment of this application; Figure 9 This is a schematic diagram of the structure of a communication device provided in an embodiment of this application; Figure 10 This is another structural schematic diagram of a communication device provided in one embodiment of this application; Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0015] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0016] In practical applications of real-time communication, the complexity of the network environment (such as fluctuations in public Wi-Fi signals, mobile network switching, cross-regional network latency, and bandwidth contention) often leads to frequent problems such as communication lag, audio-video desynchronization, and blurry images. Most related technologies rely on passive detection mechanisms, which monitor network packet loss rate, latency jitter, bandwidth fluctuations, and other indicators in real time. Only after determining that communication lag has occurred will they trigger optimization operations such as reducing video resolution, decreasing frame rate, and compressing audio bitrate.
[0017] However, such remedial measures have significant limitations: First, the response inherently involves a delay; from the occurrence of a stutter to the parameter adjustment taking effect, several seconds are often required. During this time, stuttering, screen tearing, or audio-visual distortion has already caused irreversible damage to the user experience (such as missing key information in meetings or interrupting online classroom interaction). Second, the adjustment strategy lacks foresight and precision, passively adjusting only based on the network anomalies that have already occurred, failing to anticipate network fluctuation trends (such as an impending bandwidth drop or switching to a weak signal area), leading to poor adjustment effects or over-adjustment (such as unnecessarily reducing high-definition image quality or causing audio distortion), making it difficult to balance communication quality and user experience. In addition, in multi-device concurrent communication scenarios (such as multi-person video conferencing or multi-terminal simultaneous connection), some technologies consume a lot of resources for remedial measures, which may further exacerbate network congestion, creating a vicious cycle of "stuttering-adjustment-resource consumption-more stuttering," and failing to meet the real-time communication needs of complex scenarios.
[0018] To solve or partially solve the above-mentioned technical problems, one embodiment of this application provides a communication method. For example... Figure 1 The communication method illustrated here, although showing a logical sequence of steps in the schematic diagram, may in some cases be performed in a different order than that shown in the figures. Specifically, this communication method can be applied to a sending end, which may include, but is not limited to, one or more of smartphones, tablets, laptops, and desktop computers.
[0019] The sending end can communicate with the receiving end, and this communication can include, but is not limited to, voice calls and video calls. The sending and receiving ends can be one-to-one, one-to-many, many-to-one, or many-to-many. For example, it could be two people having a one-to-one voice call, one teacher giving a one-to-many lesson to multiple students, a customer service team providing one-to-many joint Q&A to a single user, or multiple people conducting a many-to-many video conference, etc.
[0020] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.
[0021] according to Figure 1 The communication method shown includes at least steps S110 to S130, which are described in detail below: In step S110, the future Quality of Service (QoS) vector is predicted based on the current communication characteristics.
[0022] Current communication characteristics can include multi-dimensional features related to the current communication. These characteristics can reflect network transmission performance, media data attributes, and external environmental conditions, thus providing data support for predicting QoS vectors for future periods. Current communication characteristics may include, but are not limited to: network characteristics, media characteristics, and / or environmental characteristics. Network characteristics characterize the transmission performance of the communication network, directly affecting the stability and timeliness of data transmission; network characteristics may include, but are not limited to, packet loss rate, round-trip time (RTT), and / or bandwidth. Media characteristics characterize the encoding and transmission attributes of media data such as audio and video, determining the transmission efficiency and presentation quality of media data; media characteristics may include, but are not limited to, video frame rate and / or audio bitrate. Environmental characteristics characterize the external scenario and usage status of the communication process, determining dynamic changes in the communication scenario based on environmental characteristics; environmental characteristics may include, but are not limited to, mobility status, WiFi / cellular network switching, etc.
[0023] A future QoS vector can refer to a QoS vector for a future time period, such as a QoS vector for the next 0.5 seconds to 3 seconds. A QoS vector can include, but is not limited to, QoS levels, audio processing recommendations, video processing recommendations, and / or background processing recommendations. A QoS level can be a quantification of service quality; for example, QoS levels can range from 1 to 10, where a higher QoS level indicates better service quality.
[0024] In step S120, the acquired audio stream and / or video stream are processed according to the future QoS vector to obtain the target data.
[0025] When the network is working properly, in a voice call scenario, the sending end can collect the audio stream and transmit it to the receiving end; in a security monitoring scenario, the sending end may be a camera, which only transmits the video stream to the receiving end; in a video call scenario, the sending end can transmit both audio and video streams to the receiving end simultaneously.
[0026] Based on future QoS vectors, the method for processing the acquired audio and / or video streams can be determined to obtain the target data. The audio and / or video streams to be processed can be data sent in the future time period corresponding to the future QoS vector. Specifically, the duration of processing the audio and / or video streams to obtain the target data can be determined, and the acquisition time corresponding to the audio and / or video streams to be processed can be determined based on the future time period corresponding to the future QoS vector.
[0027] For example, the future time period corresponding to the future QoS vector refers to the future 0.5s to 3s. Based on big data or historical experience, it can be determined that the processing time for audio streams and / or video streams to obtain target data is 0.2s. Then, the audio streams and / or video streams collected in the future 0.3s to 2.8s can be processed to obtain target data, ensuring that the processed audio streams and / or video streams can be sent in the future 0.5s to 3s.
[0028] Optionally, since the predicted future QoS vector corresponds to a future time period that is close to the current time, the target data can also be obtained by directly processing the currently acquired audio stream and / or video stream.
[0029] In the future, when the QoS vector is different, the way audio streams and / or video streams will be processed will also be different, thus obtaining different target data; the processing methods will be detailed later.
[0030] In step S130, the target data is sent to the receiving end so that the receiving end can input the target data into the generation model and obtain the communication content.
[0031] After receiving the target data, the sending end can send it to the receiving end. Upon receiving the target data, the receiving end can input it into a generative model, which can then reconstruct all or part of the communication content based on the target data. For example, if the target data includes avatar parameters, the generative model can generate a static avatar based on those parameters; if the target data includes voiceprint features and speech text, the generative model can reconstruct the complete audio.
[0032] By adopting the technical solution of this application embodiment, the sending end can predict the future QoS vector in advance based on the current communication characteristics, and schedule the audio stream and video stream in advance according to the predicted future QoS vector to obtain the target data. When transmitting the target data based on the future network, problems such as stuttering can be avoided. The receiving end can input the target data into the generation model to obtain the communication content, thereby ensuring that the receiving end can accurately obtain the communication content and achieve effective communication. In this way, by predicting the communication quality in advance and scheduling it in advance, the user's communication experience can be guaranteed.
[0033] Based on the above technical solution, as an example, the sending end can predict the future QoS vector according to the following steps: input the current communication features into the temporal convolutional network and long short-term memory network of the dual-window prediction model respectively to obtain the sudden change features and long-term trend features; fuse the sudden change features and long-term trend features based on the cross-modal attention mechanism to obtain the fused features; and predict the future QoS vector based on the fused features.
[0034] The sending end can predict future QoS vectors based on a dual-time-window prediction model. Figure 2 This is a schematic diagram of the structure of a dual-time-window prediction model provided in an embodiment of this application; see reference. Figure 2 The dual-window prediction model can include an input module, a temporal convolutional network (TCN), a long short-term memory network (LSTM), an attention fusion module, a prediction module, and a future QoS vector output module.
[0035] In this context, TCN stands for short time window, while LSTM stands for long time window. The short time window in TCN refers to the time window constructed by TCN using causal convolution and fixed-size convolutional kernels to extract recent local correlation features from time-series data. The window length is typically determined by the kernel size, dilation coefficient, and number of network layers. For example, with a kernel size of 3 and a dilation coefficient of 1, the basic short time window length is 3, meaning it only captures temporal dependencies over three consecutive time points. The short time window of TCN makes it adept at capturing short-term, high-frequency fluctuations in time-series data, such as rapid changes in packet loss rate and bandwidth over short time scales in communication scenarios (e.g., instantaneous bandwidth jitter in WiFi networks). Through feature extraction within a local window, these short-term dynamic patterns can be accurately captured.
[0036] The long window of LSTM refers to a dynamic time window constructed by LSTM through gating mechanisms and cell states. It can overcome the limitations of short-term memory and effectively capture long-term, cross-timescale dependencies in time-series data. The window length of LSTM has no fixed limit. The long window of LSTM can effectively learn temporal correlations across multiple time steps. When predicting future QoS vectors, the long window of LSTM can deeply mine long-term patterns of communication characteristics. For example, by analyzing user movement trajectories (environmental characteristics), network handover records (environmental characteristics), and video bitrate adjustment history (media characteristics) within the past minute, the trend of future QoS vector changes can be predicted.
[0037] The sending end inputs the current communication characteristics into the temporal convolutional network and long short-term memory network of the dual-window prediction model. The temporal convolutional network can predict the sudden changes in communication characteristics, while the long short-term memory network can predict the long-term trends. Sudden changes refer to the instantaneous and abrupt characteristics of communication states that occur within a short timescale (e.g., milliseconds to seconds). Based on these characteristics, instantaneous anomalies or rapid fluctuations in the communication system can be accurately captured, providing crucial information for addressing sudden network problems in QoS prediction. Long-term trends refer to the stable and regular evolution trajectory of communication states over a longer timescale (e.g., seconds to minutes), reflecting the continuous changes in network performance, media transmission, or environmental conditions. These long-term trends can provide data support for predicting long-term service quality changes in QoS prediction.
[0038] The sudden change features and long-term trend features are input into the attention fusion module. The attention fusion module can fuse the sudden change features and long-term trend features based on a cross-modal attention mechanism to obtain fused features. The prediction module may include a fully connected layer. The fused features are input into the fully connected layer, which can predict the future QoS vector.
[0039] By adopting the technical solution of this application embodiment, the dual time window of the prediction model achieves full-dimensional coverage of short-term and long-term features. The complementarity and cross-validation of the two types of features can significantly improve the prediction accuracy of future QoS vectors and the anti-interference ability of the model, thereby adapting to the prediction needs of various dynamic communication scenarios. Furthermore, audio streams and / or video streams can be processed in advance based on the predicted future QoS vectors, thereby ensuring the user's communication experience.
[0040] Based on the above technical solution, as an example, when processing the collected audio stream and / or video stream according to the future QoS vector, the future communication mode can be determined first according to the future QoS vector, and then the collected audio stream and / or video stream can be processed according to the future communication mode to obtain the target data.
[0041] The future communication pattern represents the communication pattern for the future time period corresponding to the future QoS vector. Different communication patterns process audio and / or video streams differently, resulting in different target data and different communication content obtained by the receiver based on the generative model.
[0042] In one embodiment, four communication modes, M0, M1, M2, and M3, may be included. The M0 communication mode can be a conventional audio and video mode. In the M0 communication mode, the transmitting end does not need to process the acquired audio and / or video streams, but can directly determine the acquired audio and / or video streams as target data and transmit them. After receiving the audio and / or video streams, the receiving end does not need to input the audio and / or video streams into the generation model, but can directly determine the audio and / or video streams as communication content.
[0043] The M1 communication mode can be a bitrate / resolution reduction mode. In M1 communication mode, the sending end can reduce the bitrate of the audio stream and / or reduce the resolution of the video stream to obtain the target data and transmit the target data to the receiving end. The receiving end can directly identify the audio stream and / or video stream with reduced bitrate / resolution as the communication content, or it can indicate the bitrate / resolution of the audio stream and / or video stream to obtain the communication content.
[0044] The M2 communication mode can be an audio + dynamic avatar mode. In the M2 communication mode, the sending end can extract the voiceprint features of the audio stream and convert the audio into speech text. The sending end can extract the avatar parameters at various moments in the video stream, and use the voiceprint features, speech text and / or avatar parameters as target data, and transmit the target data to the receiving end. The receiving end can use the generative model to reconstruct the audio from the voiceprint features and speech text, and generate a dynamic avatar based on the avatar parameters, thus communicating the content.
[0045] The M3 communication mode can be a text + static avatar mode. In the M3 communication mode, the sending end can convert audio into speech text, and the sending end can extract static avatar parameters from the video stream, use the speech text and / or avatar parameters as target data, and transmit the target data to the receiving end; the receiving end can use the generated model avatar parameters to generate a static avatar, and determine the static avatar and speech text as the communication content.
[0046] In one embodiment, under different future QoS vectors, different processing can be applied to the avatars and / or background images in the audio and video streams. Different communication modes can be obtained by combining various processing methods for the avatars and / or background images in the audio and video streams. For example, processing methods for the audio stream include speech-to-text conversion, and processing methods for the video stream include reducing the resolution of the video stream, resulting in a text + reduced resolution communication mode. The four communication modes M0, M1, M2, and M3 represent only examples of some communication modes.
[0047] By adopting the technical solution of this application embodiment, different future communication modes are determined according to different future QoS vectors, thereby performing different processing on audio streams and / or video streams, resulting in different amounts of target data. This ensures that the corresponding target data can be completely transmitted under different QoS vectors, thereby ensuring that the receiving end can obtain the communication content based on the complete target data, and thus ensuring that the communication semantics can be accurately conveyed under the future QoS vector.
[0048] Based on the above technical solution, as an embodiment, when determining the future communication mode according to the future QoS vector, the communication mode that minimizes the corresponding QoS loss value can be determined from multiple communication modes under the future QoS vector. Specifically, the audio loss value, video loss value, and generation delay corresponding to each communication mode under the future QoS vector can be obtained; based on the audio loss value, video loss value, and generation delay corresponding to each communication mode, the QoS loss value corresponding to each communication mode under the future QoS vector can be determined; and the communication mode with the minimum corresponding QoS loss value among the different communication modes can be determined as the future communication mode.
[0049] The audio branch of the generative model processes the target data corresponding to the audio stream. Different communication modes result in different target data for the audio stream, and the audio branch of the generative model processes this target data differently, thus leading to different audio loss values for each branch. It's understandable that even though the audio branch is trained, its predictions may not be 100% accurate; there may still be differences between the predictions and the actual results. The audio loss value is calculated based on the difference between the predictions output by the trained audio branch and the corresponding actual results. For example, in M2 communication mode, voiceprint feature samples and speech-text samples can be obtained from real audio samples. These samples are then input into the trained audio branch to obtain reconstructed audio samples. The audio loss value for M2 communication mode can then be calculated based on the difference between the reconstructed audio samples and the real audio samples. Multiple samples from the same communication mode can be input into the trained audio branch, and the audio loss value for each sample can be calculated. The average of these audio loss values is then used as the audio loss value for that communication mode. Optionally, the audio loss value under different communication modes can be determined based on big data or historical experience.
[0050] Similarly, the video loss value corresponding to different communication modes can be determined. The video loss value is calculated based on the difference between the predicted result output by the video branch of the trained generative model and the corresponding real result. Multiple samples from the same communication mode can be input into the trained video branch, and the video loss value corresponding to each sample can be calculated. The average of the video loss values of the multiple samples is determined as the video loss value corresponding to that communication mode. Optionally, the video loss value for different communication modes can be determined based on big data or historical experience.
[0051] Under different future QoS vectors, the transmission performance of target data varies depending on the data type. For example, when a future QoS vector indicates poor future network transmission performance (such as high packet loss rate, high latency, and low bandwidth), audio stream data transmission, which is highly sensitive to real-time performance and bandwidth, is prone to problems such as stuttering, packet loss, and latency accumulation. However, for voice and text data, which have lower transmission resource requirements and stronger fault tolerance, the transmission process can usually remain smooth and stable. The data type of the target data to be transmitted can be determined based on different communication modes, and the transmission accuracy of different data types under different future QoS vectors can be determined based on big data statistics or historical experience. Transmission accuracy refers to the ratio of the amount of valid data successfully received by the receiver to the amount of data originally sent by the sender during data transmission. Based on the transmission accuracy of different data types under different future QoS vectors, the weights of different data types under different future QoS vectors can be determined.
[0052] Based on the weights of target data of different data types transmitted under different future QoS vectors, the data types of each target data transmitted under each communication mode, and the audio and video loss values under each communication mode, the corresponding audio and video loss values for each communication mode under the future QoS vector can be determined. For example, based on the accuracy of transmitting voiceprint features and speech text under a certain future QoS vector, the weight of the target data (voiceprint features and speech text) corresponding to the audio stream transmitted under that future QoS vector can be determined to be 0.5. In M2 communication mode, voiceprint features and speech text are transmitted, and the audio loss value corresponding to voiceprint features and speech text under M2 communication mode is 0.3. Therefore, the audio loss value corresponding to M2 communication mode under that future QoS vector can be determined to be 0.5 × 0.3 = 0.15.
[0053] The generation delay for each communication mode under the future QoS vector can be determined based on big data or historical experience. Generation delay refers to the time delay between the sender transmitting the target data and the receiver receiving the target data and obtaining the communication content.
[0054] Optionally, the audio loss value, video loss value, and generation delay corresponding to different communication modes under different future QoS vectors can be predetermined and stored. When it is necessary to determine the communication mode corresponding to the future QoS vector, the stored audio loss value, video loss value, and generation delay corresponding to different communication modes under the future QoS vector can be directly weighted and summed to obtain the QoS loss value corresponding to each communication mode under the future QoS vector. The QoS loss value corresponding to a certain communication mode can be determined according to the following formula: ; in, This represents the QoS loss value corresponding to the communication mode. This represents the audio loss value corresponding to this communication mode under the future QoS vector. This represents the video loss value corresponding to this communication mode under the future QoS vector. The generation delay corresponding to this communication mode under the future QoS vector; , and These are the weighting coefficients.
[0055] Based on the QoS loss values corresponding to different communication modes under the future QoS vector, the communication mode with the smallest QoS loss value can be selected as the future communication mode from multiple communication modes. The future communication mode can be determined according to the following mode selection strategy function: ; in, Representation mode selection strategy function, Characterization from , , and Choose one of the four communication modes. As a future communication mode; Characteristic of the chosen communication mode It can be made in ( In this future time period Minimum; of which The meaning of the remaining characters can be found in the previous text.
[0056] The technical solution adopted in this application embodiment selects the future communication mode that has the smallest QoS loss value under the future QoS vector, thereby ensuring the user's communication experience. In particular, it can maintain high-definition and smooth transmission when the network is good, and avoid problems such as lag and packet loss when the network is poor. It adapts to the dynamic communication needs of multiple scenarios and improves service reliability and user satisfaction.
[0057] Based on the above technical solution, as an example, to avoid frequent switching of communication modes, a hysteresis threshold mechanism can be set. This involves obtaining the current QoS vector and the hysteresis interval corresponding to the current communication mode; determining the future communication mode only when the future QoS vector is greater than the upper limit of the hysteresis interval or less than the lower limit; and maintaining the current communication mode when the future QoS vector is neither greater than the upper limit nor less than the lower limit.
[0058] A QoS vector can represent the QoS level; optionally, a higher QoS level corresponds to a higher quality of service. Different communication modes correspond to different hysteresis intervals. Due to the hysteresis threshold mechanism, the QoS level intervals corresponding to two adjacent communication modes may overlap.
[0059] For example, the QoS level range is 1~10, and the correspondence between communication modes and hysteresis intervals (QoS level intervals) includes (M0: 7~10), (M1: 5~8), (M2: 3~6), and (M3: 1~4), meaning that the QoS level interval corresponding to communication mode M0 is 7~10. If the current communication mode is M1, the current QoS level is 6.5, and the predicted future QoS vector corresponds to a future QoS level of 7.5, although 7.5 falls within the QoS level interval corresponding to communication mode M0, because 7.5 is not greater than the upper limit of the QoS level interval corresponding to communication mode M1 (8) and not less than the lower limit of the QoS level interval corresponding to communication mode M1 (5), no adjustment to the communication mode is needed, and therefore no determination of the future communication mode is required. If the predicted future QoS vector corresponds to a future QoS level of 8.5, which is greater than the upper limit of the QoS level interval corresponding to communication mode M1 (8), then the future communication mode can be determined based on the future QoS vector.
[0060] The technical solution adopted in this application can effectively avoid frequent switching of communication modes under similar QoS vectors through the hysteresis threshold mechanism, thereby reducing switching overhead and service interruption risk. On the one hand, it can prevent invalid switching caused by minor fluctuations in QoS, and on the other hand, it can adapt in a timely manner when the network status changes significantly, ensuring continuous and smooth transmission of communication services and improving the user's communication experience.
[0061] Based on the above technical solutions, as an embodiment, future communication modes may include, but are not limited to, at least: audio degradation mode, video degradation mode, and / or background degradation mode. Audio degradation modes may include, but are not limited to, audio compression mode, audio reconstruction mode, and / or summary text mode; video degradation modes may include, but are not limited to, resolution reduction mode, dynamic avatar mode, and / or static avatar mode. Audio degradation modes, video degradation modes, and / or background degradation modes can be combined to obtain other communication modes. For example, the M1 communication mode described above may be a combination of audio compression mode and resolution reduction mode; the M2 communication mode described above may be a combination of audio reconstruction mode and dynamic avatar mode.
[0062] In audio compression mode, the bit rate of the acquired audio stream can be compressed to obtain a compressed audio stream with a reduced bit rate while ensuring auditory perception quality. This compressed audio stream can be transmitted as the target data corresponding to the audio stream. The compressed audio stream can effectively adapt to complex network scenarios such as low bandwidth and high packet loss, reducing the consumption of transmission resources.
[0063] In audio reconstruction mode, voiceprint feature extraction can be performed on the acquired audio stream to obtain the extracted voiceprint features. The acquired audio stream is then processed into speech-to-text to obtain the corresponding speech-text. The voiceprint features and speech-text are transmitted as the target data corresponding to the audio stream. After receiving the voiceprint features and speech-text, the receiving end can generate reconstructed audio based on a generative model. The voiceprint features of the reconstructed audio are consistent with those extracted from the audio stream. This preserves user voiceprint recognition, ensures scenario reliability, and significantly reduces transmission volume by only transmitting voiceprint features and speech-text, thus guaranteeing accurate delivery of audio communication content even with poor network performance.
[0064] In summary text mode, semantic summary text generation can be performed on the acquired audio stream. First, the acquired audio stream is converted from speech to text to obtain speech text. Then, the speech text is summarized and generalized to obtain semantic summary text. The semantic summary text is then transmitted as the target data corresponding to the audio stream.
[0065] In the resolution reduction mode, the acquired video stream can be processed to reduce its resolution. By adjusting the pixel dimension through an adaptive downsampling algorithm, a low-resolution video stream is generated while balancing visual perception quality. This low-resolution video stream is then transmitted as the target data corresponding to the video stream, thereby effectively reducing transmission bandwidth usage and storage overhead.
[0066] In dynamic avatar mode, avatar parameters can be extracted in real time from the captured video stream, yielding avatar parameters at multiple moments. This involves extracting multiple video frames containing the avatar, locating the avatar region within each frame using a face detection algorithm, and then using a feature extraction model to extract avatar parameters such as pose angle, facial key point coordinates, and / or facial expression features from each frame, forming a temporal sequence of avatar parameters. This sequence includes avatar parameters from multiple moments. These multiple moments of avatar parameters are then transmitted as the target data corresponding to the video stream. By transmitting dynamic avatar parameters, the receiving end can generate a dynamic avatar based on a generative model. The mouth and facial expressions of this dynamic avatar can change in real time following the speaker's mouth and facial expressions, accurately conveying the speaker's emotions and state. Furthermore, transmitting multiple moments of avatar parameters instead of the complete video stream significantly reduces bandwidth consumption and transmission latency, adapting to complex network scenarios with low bandwidth and high packet loss, and solving the problem of stuttering during high-definition video transmission.
[0067] In static avatar mode, static avatar parameters can be extracted from avatars in the captured video stream to obtain the static avatar parameters corresponding to a specific video frame. Face detection algorithms can be used to locate the avatar region within a video frame, and then a feature extraction model can be used to extract static avatar parameters such as pose angle, facial key point coordinates, and / or facial expression features. These static avatar parameters are then transmitted as target data corresponding to the video stream. Specifically, when a speaker is detected as about to speak, the speaker's avatar parameters can be extracted to allow the user to identify the speaker.
[0068] In background degradation mode, background images can be extracted from the captured video stream, and background tags can be generated based on the semantics of the background images. These background tags are then transmitted as target data corresponding to the video stream. For example, if the background is a study, the background tag "study" can be generated. After receiving the "study" tag, the receiving end can generate a study scene based on the generation model and use the generated study scene as the video background. In this way, only the background tag needs to be transmitted to reconstruct the video background, which can greatly save transmission resources and has minimal impact on the transmission of communication content.
[0069] Based on the above technical solutions, and considering security and confidentiality issues, digital watermarks can be added to the transmitted target data. Specifically, a session identifier (Session ID) can be obtained, and the session identifier and time slice information can be converted into multiple modal digital watermarks. These multiple modal digital watermarks are then added to the target data, each with the same modality as the digital watermark, resulting in target data with added digital watermarks. This allows the receiving end to perform cross-modal verification based on the digital watermarks extracted from the target data.
[0070] A session identifier is a globally unique identifier assigned to distinguish a single continuous communication process (such as a voice call, video conference, or data transmission session). Data management and state tracking based on session identifiers ensure that various types of data within the same session maintain correlation and consistency during transmission, processing, and storage. Time-slice information refers to the information used to characterize individual time segments after dividing continuous communication time-series data (such as video streams and audio streams) according to preset time granularities (such as 10ms, 50ms, and 100ms). Time-slice information can include the start and end times of the time slice; based on time-slice information, segmented processing, feature extraction, and precise correlation of time-series data can be performed.
[0071] The sending end can obtain the session identifier and time slice information corresponding to the audio stream and / or video stream, and convert the session identifier and time slice information into multiple modal digital watermarks according to a preset algorithm. The modality of the digital watermark corresponds to the modality of the target data. Based on the audio stream and / or video stream corresponding to the digital watermark, and the target data corresponding to each audio stream and / or video stream, the target data corresponding to the digital watermark can be determined. The digital watermarks of each modality are added to the target data with the same modality as the digital watermark (this target data corresponds to the digital watermark), resulting in target data with added digital watermarks. For example, if the target data includes voice text and video streams, the modality of the digital watermark can include text modality and video modality. The text modality digital watermark can be added to the voice text to obtain voice text with added digital watermarks, and the video modality digital watermark can be added to the video stream to obtain video streams with added digital watermarks.
[0072] The sending end can send multiple target data sets with the same digital watermark (but different modalities) to the receiving end. Upon receiving the multiple target data sets, the receiving end can extract each digital watermark from each set and verify their consistency, thus performing cross-modal verification. If the digital watermarks of the multiple target data sets are consistent, the cross-modal verification passes, proving secure transmission, and the communication content can be obtained from the multiple target data sets. If the digital watermarks of the multiple target data sets are inconsistent, the cross-modal verification fails, indicating a potential transmission problem. Therefore, the receiving end can proactively downgrade the communication mode to a preset communication mode. The preset communication mode refers to the communication mode corresponding to the lowest (below a threshold) future QoS vector. For example, among the four communication modes M0 to M3, mode M3 can be designated as the preset communication mode.
[0073] Digital watermarks can be embedded in the target data for both the audio and video modalities using the following formula: ; ; in, For the target data of audio modalities without digital watermarks, For the target data of the audio modality with added digital watermark, For target data of video modalities without digital watermarks, For the target data of the video modality with added digital watermark, For digital watermarks generated based on session identifiers and time slice information, and These are the modality transfer functions for digital watermarking. Digital watermarks can be converted into audio modalities. Digital watermarks can be converted into video modalities. and The preset watermark strength coefficient can be used to control the strength of the digital watermark embedding.
[0074] The receiving end can extract the digital watermark to be verified from the target data of the audio modality and convert the digital watermark to a preset modality. The process involves extracting the digital watermark to be verified from the target data of the video modality and converting the digital watermark to a preset modality. The receiving end can verify... and Whether they are consistent determines whether the cross-modal validation has passed; in and When consistent, cross-modal validation is considered passed; and If there is an inconsistency, the cross-modal validation is deemed unsuccessful.
[0075] The technical solution adopted in this application generates a multimodal digital watermark by using session identifiers and time slice information and embeds it into the target data across modalities. This enables data traceability by using session identifiers and time slice information, and strengthens anti-counterfeiting capabilities through cross-modal verification, effectively preventing the risks of data tampering, theft, and forgery. At the same time, it adapts to the security requirements of communication scenarios, ensuring data confidentiality and integrity without affecting transmission efficiency and service experience, thereby improving the security protection level of the communication system.
[0076] Based on the above technical solution, when the communication scenario is a multi-person conference and the sending end is a participant in the multi-terminal conference, the priority of multiple participants can be determined separately, and the communication mode can be determined in combination with the priority of the participants to ensure the audio and video transmission quality of high-priority participants.
[0077] Therefore, when the sending end processes the collected audio stream and / or video stream according to the future QoS vector, it can obtain its own first priority and determine the first future communication mode according to the future QoS vector and the first priority; it processes the collected audio stream and / or video stream according to the first future communication mode to obtain the target data and sends the target data to the receiving end.
[0078] Optionally, the priority of multiple participants can be determined based on their roles (moderator, speaker, regular participant), and a pre-defined correspondence between different roles and priorities can be established. Since a participant's role may change, their priority may also change accordingly. For example, if a participant was a regular participant in the previous moment and is now a speaker, their priority may change. Optionally, the priority of multiple participants can be pre-set according to actual needs.
[0079] When determining the future communication mode, the sending end can refer to the method described above. Based on the future QoS vector, an initial future communication mode is first determined. Then, the initial future communication mode is dynamically adjusted according to the first priority to obtain the first future communication mode. For example, if the first priority is higher than the upper priority limit, the initial future communication mode can be increased by one level to obtain the first future communication mode; for example, the M3 communication mode can be increased to the M2 communication mode. Conversely, if the first priority is lower than the lower priority limit, the initial future communication mode can be decreased by one level to obtain the first future communication mode; for example, the M2 communication mode can be decreased to the M3 communication mode. The lower priority limit is lower than the upper priority limit, and both the upper and lower priority limits can be set according to actual needs.
[0080] The first future communication mode is one type of communication mode. After determining the first future communication mode, the sending end can process the audio stream and / or video stream according to the first future communication mode to obtain the target data, and then send the target data to the receiving end.
[0081] By adopting the technical solution of this application embodiment, the sending end determines the first future communication mode by its own first priority and the future QoS vector, which can adapt to multi-terminal conference scenarios and give priority to ensuring the audio and video transmission quality of high-priority participants; it avoids the waste of network resources and can allocate resources differently according to the importance of the participants, balance the multi-terminal experience in complex network environments, and improve the stability and relevance of conference communication.
[0082] In one embodiment, such as Figure 3 As shown, a communication method is provided. Although the logical order is illustrated in the step diagram, in some cases, the steps shown or described may be performed in a different order than that shown in the figures. Specifically, this communication method can be applied to a receiving end, wherein the sending end may be one or more of a smartphone, tablet computer, laptop computer, and desktop computer, including but not limited to. These will be described in detail below. It should be noted that the order of description of the following embodiments is not intended to limit the priority of the embodiments.
[0083] according to Figure 3 The communication method shown includes at least steps S310 to S320, which are described in detail below: In step S310, the target data sent by the sending end is received.
[0084] In step S320, the target data is input into the generation model to obtain the communication content.
[0085] The target data is obtained by the sending end predicting the future QoS vector based on the current communication characteristics, and then processing the collected audio stream and / or video stream according to the future QoS vector.
[0086] Current communication characteristics can include multi-dimensional characteristics related to the current communication, including but not limited to: network characteristics, media characteristics, and / or environmental characteristics. Future QoS vectors can refer to QoS vectors for future periods, and QoS vectors can include but are not limited to: QoS levels, audio processing recommendations, video processing recommendations, and / or background processing recommendations.
[0087] The sending end can acquire current communication features and input them into the temporal convolutional network and long short-term memory network of the dual-window prediction model to obtain abrupt change features and long-term trend features. Based on a cross-modal attention mechanism, the abrupt change features and long-term trend features are fused to obtain fused features. Based on the fused features, the future QoS vector is predicted. The sending end can determine the future communication mode based on the future QoS vector; according to the future communication mode, the acquired audio stream and / or video stream are processed to obtain target data, which is then sent to the receiving end. The specific method for the sending end to determine the target data can be referred to the previous text and will not be repeated here.
[0088] After receiving the target data, the receiving end can input the target data into the generative model. The generative model can then reconstruct all or part of the communication content based on the target data. The generative model is a multimodal model, capable of extracting communication content for target data of different modalities. For example, when the target data includes avatar parameters, the generative model can generate a static avatar based on those parameters; when the target data includes voiceprint features and speech text, the generative model can reconstruct the complete audio.
[0089] By adopting the technical solution of this application embodiment, the sending end can predict the future QoS vector in advance based on the current communication characteristics, and schedule the audio stream and video stream in advance according to the predicted future QoS vector to obtain the target data. When transmitting the target data based on the future network, problems such as stuttering can be avoided. The receiving end can input the target data into the generation model to obtain the communication content, thereby ensuring that the receiving end can accurately obtain the communication content and achieve effective communication. In this way, by predicting the communication quality in advance and scheduling it in advance, the user's communication experience can be guaranteed.
[0090] Based on the above technical solutions, a multimodal generation model can include multiple branches that process data of different modalities. For example, a multimodal generation model can include an audio branch, a video branch, and a background branch. The audio branch can include an audio generation network, the video branch can include an avatar generation network, and the background branch can include a background generation network. Figure 4 This is a schematic diagram of the architecture of a multimodal generation model provided in an embodiment of this application; see reference. Figure 4 Audio generation networks can include, but are not limited to, Low-Rank Adaptation (LoRA) models and / or Text-to-Speech (TTS) models. LoRA is a lightweight fine-tuning technique for pre-trained large models. It inserts a low-rank matrix (low-dimensional projection space) into the pre-trained model, training only the parameters of this low-rank matrix instead of the entire model parameters, achieving efficient model adaptation and functional expansion while keeping the original parameters of the pre-trained model frozen. TTS is a technique that automatically converts text information into natural speech signals, including text preprocessing (word segmentation, syntactic analysis, prosodic annotation), acoustic modeling (mapping text to acoustic features, such as Mel spectrum), vocoder synthesis (converting acoustic features into playable speech waveforms), and finally outputting synthesized speech with intonation and prosody consistent with human speech. Avatar generation networks can include, but are not limited to, Generative Adversarial Networks (GANs) and / or Neural Radiance Fields (NeRFs). GANs can generate realistic data through adversarial training between the generator and discriminator. NeRF can be trained to learn the radiation field (including color and density information) of a scene by taking multi-view 2D images and camera parameters as input, and can generate high-fidelity 3D rendering results from any viewpoint. Background generation networks can include, but are not limited to, diffusion models, which are generative artificial intelligence models based on an iterative process of progressive noise addition and denoising.
[0091] Target data may include, but is not limited to: speech text, voiceprint features, avatar parameters at multiple time points, static avatar parameters, and / or background tags. The method for the sending end to obtain target data can be referred to the previous section. It is understood that audio and video streams can also be used as target data, but when the receiving end receives audio and / or video streams, it can directly play them without inputting them into the generative model for processing.
[0092] When the receiving end receives the speech text and voiceprint features, it can input these features and the speech text into the audio generation network of the generative model. The audio generation network can parse the semantics and prosody of the speech text through the text encoder, and extract personalized timbre features through the voiceprint features. The audio generation network can use the predicted audio prosody as a constraint, fusing the semantic information of the speech text with the personalized attributes of the voiceprint features, converting the acoustic features into playable reconstructed audio, ensuring that the reconstructed audio not only restores the semantics of the text but also perfectly matches the voiceprint features of the original speaker.
[0093] When the receiving end receives avatar parameters from multiple moments, it can input these parameters into the avatar generation network of the generative model. The avatar generation network can capture the changing patterns of the avatar parameters at different moments, ensuring the continuity of the avatar's movements and the naturalness of its expressions at adjacent moments, and finally outputting a dynamic avatar that is completely synchronized with the avatar's movements and expressions in the original video stream.
[0094] When the receiving end receives static avatar parameters, it can input the static avatar parameters into the avatar generation network of the generation model. The avatar generation network can then reconstruct a static avatar image based on the static avatar parameters.
[0095] When the receiving end receives a background label, it can input the background label into the background generation network of the generation model. The background generation network can parse the semantics of the background label and, combined with a pre-trained scene feature library, generate a background image that matches the description of the background label.
[0096] By reconstructing audio and accurately matching the original voiceprint and semantics, the reliability of voice interaction and identity verification scenarios can be guaranteed; dynamic avatars can restore the continuity of actions and expressions in sequence, while static avatars can maintain the fidelity of details and improve the naturalness of visual interaction; background images are matched with scene tags as needed to adapt to the needs of multiple scenarios.
[0097] The communication content may include, but is not limited to: reconstructed audio, animated avatars, static avatars, background images, audio streams, video streams, compressed audio streams, low-resolution video streams, speech-to-text and / or semantic summary text. Based on the communication content, the receiving end can obtain information about the communication interaction.
[0098] By employing the technical solution of this application embodiment, the sending end and receiving end can transmit target data with a small data volume, and the receiving end can obtain the communication content based on the target data, significantly reducing bandwidth consumption and transmission latency, adapting to low-bandwidth, high-packet-loss networks, thereby ensuring communication under weak network conditions. Furthermore, each generation stage is independently controllable, allowing for flexible combination of outputs, balancing real-time performance and personalized needs, and comprehensively optimizing transmission efficiency and user experience in scenarios such as multi-terminal conferencing and remote interaction.
[0099] Based on the above technical solutions, in order to ensure the consistency between audio and video modalities and avoid the situation where the lip movements of the avatars do not correspond to the audio, when training the generative model, the audio generation network and the avatar generation network of the generative model can be trained together based on the speech-lip movement consistency loss function, thereby ensuring that the audio generated by the audio generation network and the lip movements generated by the avatar generation network correspond to each other.
[0100] The training steps for the generative model may include at least the following: obtaining the audio features output by the audio generation network to be trained; obtaining the lip features corresponding to the audio features output by the avatar generation network to be trained; establishing a speech-lip consistency loss function based on the difference between the audio features and the lip features; and training the generative model to be trained based on the speech-lip consistency loss function to obtain the trained generative model.
[0101] A large training dataset can be pre-acquired, containing pairs of original audio files and corresponding video frames. Speech text, voiceprint features, and avatar parameters are extracted from the paired original audio and video frames, respectively. The speech text and voiceprint features from the training data are input into the audio generation network to be trained, resulting in reconstructed audio samples. The reconstructed audio features are then extracted, directly reflecting the core attributes of the reconstructed audio samples. The avatar parameters corresponding to these speech text and voiceprint features are input into the avatar generation network to be trained, resulting in avatar images generated by the avatar generation network. The lip-sync features of the avatar images are then extracted.
[0102] Based on the temporal synchronization of audio features and lip-sync features, the differences between them in the feature space are calculated. These differences can include, but are not limited to, temporal alignment differences, semantic association differences, and / or feature distribution differences. Temporal alignment differences can be obtained by calculating the deviation between audio prosodic rhythm and lip-sync rhythm using a dynamic time warping algorithm; semantic association differences can be obtained by utilizing attention mechanisms to mine the matching degree between audio semantic features and lip-sync action features; and feature distribution differences can be obtained by calculating the distribution distance between the two in a high-dimensional space using cosine similarity. The weighted summation of temporal alignment differences, semantic association differences, and / or feature distribution differences yields the difference between audio features and lip-sync features.
[0103] The speech-lip consistency loss function can be determined using the following formula: ; in, Let be the speech-lip-sync consistency loss function. The length of the training data, For time step index, For the first Audio characteristics of a moment For the first Lip shape characteristics at any moment This is the encoding function for audio features, used to map audio features into high-dimensional feature vectors. This is the encoding function for lip shape features, used to map lip shape features into high-dimensional feature vectors. This represents the square of the L2 norm, used to calculate the distance between two eigenvectors.
[0104] After determining the speech-lip-sync consistency loss function, other loss functions can be combined to train the generative model until the loss function converges, resulting in a well-trained generative model. The trained generative model's audio generation network and avatar generation network can achieve accurate matching between audio generation and lip-sync, meeting the needs of real-time interactive scenarios.
[0105] The technical solution adopted in this application embodiment can constrain the temporal, semantic and distributional consistency of audio features and lip features based on the speech-lip consistency loss function, so that the audio generated by the trained model is accurately synchronized with the lip movements, avoiding sound-shape misalignment. At the same time, it strengthens the dual-network collaborative adaptation capability, improves the naturalness and realism of the synthesized content, and ensures the coordination of the user's visual and auditory experience.
[0106] Based on the above technical solution, and considering security and confidentiality issues, the target data sent by the sending end can carry a digital watermark. After receiving the target data with the digital watermark added, the receiving end can extract the digital watermarks of multiple modalities of the target data; perform cross-modal verification based on the digital watermarks of the target data; and when the cross-modal verification passes, input the target data into the generation model to obtain the communication content. The digital watermarks in the target data are obtained by the sending end by converting multiple modalities of digital watermarks based on the session identifier and time slice information corresponding to the target data, and then adding the multiple modalities of digital watermarks to the target data with the same modality as the digital watermark.
[0107] The method for the sending end to add a digital watermark to the target data can be referred to the previous text. After receiving multiple target data sets with added digital watermarks, the receiving end can extract each digital watermark from the multiple target data sets and verify whether the digital watermarks are consistent, thereby performing cross-modal verification on the multiple target data sets.
[0108] If the digital watermarks of multiple target data are consistent, the cross-modal verification passes, proving that the transmission is secure, and the communication content can be obtained based on the multiple target data. If the digital watermarks of multiple target data are inconsistent, the cross-modal verification fails, proving that there may be a problem with the transmission. Therefore, the receiving end can actively downgrade the communication mode to a preset communication mode. The preset communication mode refers to the communication mode corresponding to the lowest (below the threshold) future QoS vector. For example, among the four communication modes M0 to M3, the M3 communication mode can be determined as the preset communication mode.
[0109] Figure 5 This is a schematic diagram of the digital watermarking processing flow provided in an embodiment of this application; see reference. Figure 5The sending end can generate multimodal digital watermarks based on session identifiers and time slice information, and add these watermarks to target data of different modalities. The sending end can package multimodal target data with the same digital watermark to obtain an encrypted multimodal stream, and send this encrypted multimodal stream to the receiving end. The receiving end can receive the encrypted multimodal stream and separate it to obtain target data of different modalities. The receiving end can extract the digital watermark to be verified from the audio modal target data and convert it to a preset modality, and extract the digital watermark to be verified from the video modal target data and convert it to a preset modality. The receiving end can determine whether cross-modal verification has passed by verifying whether the two digital watermarks of the preset modality are consistent; if the digital watermarks are consistent, the cross-modal verification is considered successful, and the target data is input into the generation model to obtain the communication content; if the digital watermarks are inconsistent, the cross-modal verification is considered unsuccessful, and the communication mode is downgraded.
[0110] The technical solution of this application embodiment uses digital watermarking for cross-modal verification, which can realize data traceability and enhance anti-counterfeiting capabilities, effectively preventing the risks of data tampering, theft and forgery; at the same time, it adapts to the security requirements of communication scenarios, and under the premise of ensuring data confidentiality and integrity, it does not affect transmission efficiency and business experience, thereby improving the security protection level of communication systems.
[0111] Based on the above technical solution, when the communication scenario is a multi-person conference and the receiving end is a participant in a multi-terminal conference, the priority of multiple participants can be determined separately, and the communication mode can be determined in combination with the priority of the participants to ensure the audio and video transmission quality of high-priority participants.
[0112] The receiving end can predict its own future QoS vector and obtain its own second priority; based on its own future QoS vector and second priority, it determines a second future communication mode; according to the second future communication mode, it selects the received target data; and inputs the selected target data into the generation model to obtain the communication content.
[0113] Optionally, the priority of multiple participants can be determined based on their roles or can be pre-set according to actual needs.
[0114] The method for the receiver to predict its own future QoS vector can be similar to that for the sender. The receiver can determine an initial future communication mode based on its own future QoS vector, and then dynamically adjust this initial future communication mode according to a second priority to obtain a second future communication mode. For example, if the second priority is higher than the upper priority limit, the initial future communication mode can be increased by one level to obtain the second future communication mode; if the second priority is lower than the lower priority limit, the initial future communication mode can be decreased by one level to obtain the second future communication mode.
[0115] After determining the second future communication mode, the receiving end can select the received target data and then input the selected target data into the generation model to obtain the communication content. For example, when the first future communication mode is M2 and the second future communication mode is M3, the target data transmitted by the sending end based on the M2 communication mode includes voice text, voiceprint features, and avatar parameters at multiple times. Then, the receiving end can select voice text and one avatar parameter at a specific time from the target data according to the M3 communication mode, and generate a static avatar based on the avatar parameter at that specific time. The voice text and the static avatar are then directly displayed, thereby realizing the M3 communication mode at the receiving end.
[0116] It is understandable that if the second future communication mode is superior to the first future communication mode, but the target data can only support the first future communication mode and not the second future communication mode, then even if the receiving end determines the second future communication mode, it can still only process the target data based on the first future communication mode to obtain the communication content.
[0117] By adopting the technical solution of this application embodiment, the receiving end determines the second future communication mode through its own second priority and future QoS vector, which can adapt to multi-terminal conference scenarios and prioritize the audio and video transmission quality of high-priority participants; it avoids the waste of network resources and can allocate resources differently according to the importance of participants, balance the multi-terminal experience in complex network environments, and improve the stability and relevance of conference communication.
[0118] Figure 6 This is a timing diagram of a communication method provided in an embodiment of this application; see reference. Figure 6 The following steps can be performed sequentially from t0 to t5: Time t0: Predict future QoS vectors; At time t1: Determine the future communication mode from multiple communication modes based on the future QoS vector; At time t2: The audio and video streams are processed according to the future communication model to obtain the target data; At time t3: Embed a digital watermark in the target data and send the target data with the embedded digital watermark. At time t4: The receiver verifies the consistency of the digital watermark in the target data across modalities; At time t5: If the cross-modal validation passes, the target data is input into the multimodal generation model and the communication content is output; if the cross-modal validation fails, the communication mode is downgraded to the M3 communication mode.
[0119] To facilitate better implementation of the communication method of this application, this application also provides a communication system based on the above-described communication method. The meanings of the terms used are the same as in the communication method described above, and specific implementation details can be found in the descriptions of the method embodiments. Please refer to... Figure 7 , Figure 7 This is a schematic diagram of the structure of a communication system provided in an embodiment of this application, wherein the communication system includes: The sending end is used to predict the future QoS vector based on the current communication characteristics; based on the future QoS vector, it processes the collected audio stream and / or video stream to obtain the target data, and sends the target data to the receiving end; The receiving end is used to input the target data into the generative model to obtain the communication content.
[0120] In one embodiment, the current communication characteristics include: network characteristics, media characteristics, and / or environmental characteristics; The sending end predicts the future QoS vector based on current communication characteristics, including: The transmitting end inputs the current communication features into the temporal convolutional network and long short-term memory network of the dual-window prediction model to obtain sudden change features and long-term trend features. The sudden change features and the long-term trend features are fused based on a cross-modal attention mechanism to obtain fused features; Based on the fused features, the future QoS vector is predicted.
[0121] In one embodiment, the transmitting end processes the acquired audio and / or video streams according to a future QoS vector to obtain target data, including: The sending end determines the future communication mode based on the future QoS vector; The transmitting end processes the acquired audio stream and / or video stream according to the future communication mode to obtain the target data.
[0122] In one embodiment, the sending end determines the future communication mode based on the future QoS vector, including: The transmitting end obtains the audio loss value, video loss value, and generation delay corresponding to different communication modes under the future QoS vector; Based on the audio loss value, the video loss value, and the generation delay corresponding to each of the different communication modes, determine the QoS loss value corresponding to each of the different communication modes under the future QoS vector; The communication mode with the smallest QoS loss value among different communication modes is determined as the future communication mode.
[0123] In one embodiment, the sending end determines the future communication mode based on the future QoS vector, including: The sending end obtains the current QoS vector and the hysteresis interval corresponding to the current communication mode; The future communication mode is determined when the future QoS vector is greater than the upper limit of the hysteresis interval or less than the lower limit of the hysteresis interval.
[0124] In one embodiment, the future communication mode includes at least: an audio degradation mode, a video degradation mode, and / or a background degradation mode; The transmitting end processes the acquired audio stream and / or video stream according to the future communication mode, including one or more of the following steps: In the audio degradation mode, the acquired audio stream is subjected to compression processing, voiceprint feature extraction processing, speech-to-text processing, and / or semantic summary text generation processing. In the video degradation mode, the acquired video stream is subjected to resolution reduction processing, extraction of avatar parameters at multiple times, and / or extraction of static avatar parameters. In the background degradation mode, a background image is extracted from the acquired video stream, and a background label is generated based on the background image.
[0125] In one embodiment, the generative model includes: an audio generation network, an avatar generation network, and a background generation network; the target data includes: speech text, voiceprint features, avatar parameters at multiple time points, static avatar parameters, and / or background labels; The receiving end inputs the target data into the generation model to obtain the communication content, including one or more of the following steps: The receiving end inputs the speech text and the voiceprint features into the audio generation network to predict the audio prosody, and generates reconstructed audio based on the audio prosody, the text, and the voiceprint features; The avatar parameters at multiple times are input into the avatar generation network to generate dynamic avatars; The static avatar parameters are input into the avatar generation network to generate a static avatar. The background label is input into the background generation network to generate a background image.
[0126] In one embodiment, the generative model includes: an audio generation network and an avatar generation network; the training steps of the generative model include: Obtain the audio features output by the audio generation network to be trained; Obtain the lip features corresponding to the audio features output by the head-image generation network to be trained; Based on the differences between the audio features and the lip-sync features, a speech-lip-sync consistency loss function is established; The generative model to be trained is trained according to the speech-lip consistency loss function to obtain the trained generative model.
[0127] In one embodiment, the sending end is configured to acquire the session identifier and time slice information corresponding to the target data, and convert the session identifier and time slice information into multiple modal digital watermarks; and add the multiple modal digital watermarks to the target data with the same modality as the digital watermarks, to obtain the target data with added digital watermarks; The receiving end inputs the target data into the generation model to obtain the communication content, including: After receiving the target data with added digital watermarks corresponding to the same time slice, the receiving end extracts the digital watermarks of the target data in multiple modalities. Cross-modal verification is performed based on the digital watermark of the target data; When the cross-modal validation passes, the target data is input into the generative model to obtain the communication content.
[0128] In one embodiment, the sending end and the receiving end are participants in a multi-terminal conference; The transmitting end processes the acquired audio and / or video streams according to the future QoS vector to obtain target data, including: The sending end obtains its own first priority and determines a first future communication mode based on the future QoS vector and the first priority; The transmitting end processes the acquired audio stream and / or video stream according to the first future communication mode to obtain the target data; The receiving end inputs the target data into the generation model to obtain the communication content, including: The receiving end predicts its own future QoS vector and obtains its own second priority; The receiving end determines the second future communication mode based on its own future QoS vector and the second priority; The receiving end selects the received target data according to the second future communication mode, and inputs the selected target data into the generation model to obtain the communication content.
[0129] Figure 8 This is another structural schematic diagram of a communication system provided in an embodiment of this application; see reference. Figure 8 The communication system may include a transmitter and a receiver. The transmitter may include, but is not limited to: an audio acquisition module, a video acquisition module, a network prediction module, an intelligent encoding module, an AIGC (Artificial Intelligence Generated Content) control flow generation module, a security module, and / or a communication module. The receiver may include, but is not limited to: a communication module, an intelligent decoding module, a security restoration module, a multimodal generation module, and / or an output rendering module. The audio and video acquisition modules can acquire audio and video streams respectively. The network prediction module may include a dual-time-window prediction model, which can predict future QoS vectors. The intelligent encoding module can encode the audio and / or video streams. The AIGC control flow generation module can process the encoded audio and / or video streams to obtain target data. The security module can generate a multimodal digital watermark and embed the watermark into the target data. The communication module of the transmitter can send the target data with the embedded digital watermark to the receiver. The communication module of the receiver can receive the target data. The intelligent decoding module can decode the multimodal target data. The security restoration module can extract the digital watermark from the target data and perform cross-modal verification. The multimodal generation module can include a multimodal generation model, which can generate communication content based on target data. The output rendering module can play and display the communication content.
[0130] By adopting the technical solution of this application embodiment, the sending end can predict the future QoS vector in advance based on the current communication characteristics, and schedule the audio stream and video stream in advance according to the predicted future QoS vector to obtain the target data. When transmitting the target data based on the future network, problems such as stuttering can be avoided. The receiving end can input the target data into the generation model to obtain the communication content, thereby ensuring that the receiving end can accurately obtain the communication content and achieve effective communication. In this way, by predicting the communication quality in advance and scheduling it in advance, the user's communication experience can be guaranteed.
[0131] To facilitate better implementation of the communication method of this application, this application also provides a communication device based on the above-described communication method. The meanings of the terms used are the same as in the above-described communication method, and specific implementation details can be found in the descriptions of the method embodiments.
[0132] Please see Figure 9 , Figure 9This is a schematic diagram of the structure of a communication device provided in an embodiment of this application, wherein the communication device is applied to a transmitting end and may include: The prediction module 901 is used to predict future QoS vectors based on current communication characteristics; The processing module 902 is used to process the acquired audio stream and / or video stream according to the future QoS vector to obtain the target data; The sending module 903 is used to send the target data to the receiving end, so that the receiving end can input the target data into the generation model to obtain the communication content.
[0133] In one embodiment, the current communication characteristics include: network characteristics, media characteristics, and / or environmental characteristics; the prediction module 901 is specifically used to perform: The current communication features are input into the temporal convolutional network and long short-term memory network of the dual-window prediction model to obtain sudden change features and long-term trend features. The sudden change features and the long-term trend features are fused based on a cross-modal attention mechanism to obtain fused features; Based on the fused features, the future QoS vector is predicted.
[0134] In one embodiment, the processing module 902 is specifically used to perform: Determine the future communication mode based on the future QoS vector; According to the future communication mode, the acquired audio stream and / or video stream are processed to obtain the target data.
[0135] In one embodiment, determining the future communication pattern based on the future QoS vector includes: Under the future QoS vector, obtain the audio loss value, video loss value and generation delay corresponding to different communication modes; Based on the audio loss value, the video loss value, and the generation delay corresponding to each of the different communication modes, determine the QoS loss value corresponding to each of the different communication modes under the future QoS vector; The communication mode with the smallest QoS loss value among different communication modes is determined as the future communication mode.
[0136] In one embodiment, determining the future communication pattern based on the future QoS vector includes: Obtain the current QoS vector and the hysteresis interval corresponding to the current communication mode; The future communication mode is determined when the future QoS vector is greater than the upper limit of the hysteresis interval or less than the lower limit of the hysteresis interval.
[0137] In one embodiment, the future communication mode includes at least: an audio degradation mode, a video degradation mode, and / or a background degradation mode; The processing of the acquired audio stream and / or video stream according to the future communication mode includes one or more of the following steps: In the audio degradation mode, the acquired audio stream is subjected to compression processing, voiceprint feature extraction processing, speech-to-text processing, and / or semantic summary text generation processing. In the video degradation mode, the acquired video stream is subjected to resolution reduction processing, extraction of avatar parameters at multiple times, and / or extraction of static avatar parameters. In the background degradation mode, a background image is extracted from the acquired video stream, and a background label is generated based on the background image.
[0138] In one embodiment, the device further includes: The watermark acquisition module is used to acquire the session identifier and time slice information corresponding to the target data, and convert the session identifier and time slice information into digital watermarks of multiple modalities. A watermarking module is used to add multiple modal digital watermarks to target data of the same modality as the digital watermarks, thereby obtaining target data with added digital watermarks, so that the receiving end can perform cross-modal verification based on the digital watermarks extracted from the target data.
[0139] In one embodiment, the sending end is a participant in a multi-terminal conference; the processing module 902 is specifically used to execute: Obtain its own first priority, and determine a first future communication mode based on the future QoS vector and the first priority; According to the first future communication mode, the collected audio stream and / or video stream are processed to obtain the target data.
[0140] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of a communication device provided in an embodiment of this application, wherein the communication device is applied at a receiving end and may include: The receiving module 101 is used to receive target data sent by the sending end; Input module 102 is used to input the target data into the generation model to obtain communication content; The target data is obtained by the sending end predicting the future QoS vector based on the current communication characteristics, and processing the collected audio stream and / or video stream according to the future QoS vector.
[0141] In one embodiment, the generative model includes: an audio generation network, an avatar generation network, and a background generation network; the target data includes: speech text, voiceprint features, avatar parameters at multiple time points, static avatar parameters, and / or background labels; the input module 102 is specifically used to perform one or more of the following steps: The audio generation network is input with the spoken text and the voiceprint features to predict the audio prosody, and then generates reconstructed audio based on the audio prosody, the spoken text, and the voiceprint features. The avatar parameters at multiple times are input into the avatar generation network to generate dynamic avatars; The static avatar parameters are input into the avatar generation network to generate a static avatar. The background label is input into the background generation network to generate a background image.
[0142] In one embodiment, the generative model includes: an audio generation network and an avatar generation network; the training steps of the generative model include: Obtain the audio features output by the audio generation network to be trained; Obtain the lip features corresponding to the audio features output by the head-image generation network to be trained; Based on the differences between the audio features and the lip-sync features, a speech-lip-sync consistency loss function is established; The generative model to be trained is trained according to the speech-lip consistency loss function to obtain the trained generative model.
[0143] In one embodiment, the input module 102 is specifically used to perform: After receiving the target data with added digital watermark, extract the digital watermarks of the target data in multiple modalities; Cross-modal verification is performed based on the digital watermark of the target data; When the cross-modal validation passes, the target data is input into the generative model to obtain the communication content; The digital watermark in the target data is obtained by the sending end from converting multiple modal digital watermarks based on the session identifier and time slice information corresponding to the target data, and then adding the multiple modal digital watermarks to the target data with the same modality as the digital watermark.
[0144] In one embodiment, the receiving end is a participant in a multi-terminal conference; the input module 102 is specifically used to perform: Predict its own future QoS vector and obtain its own second priority; Based on its own future QoS vector and the second priority, a second future communication mode is determined; According to the second future communication mode, the received target data is selected; The selected target data is input into the generation model to obtain the communication content.
[0145] By adopting the technical solution of this application embodiment, the sending end can predict the future QoS vector in advance based on the current communication characteristics, and schedule the audio stream and video stream in advance according to the predicted future QoS vector to obtain the target data. When transmitting the target data based on the future network, problems such as stuttering can be avoided. The receiving end can input the target data into the generation model to obtain the communication content, thereby ensuring that the receiving end can accurately obtain the communication content and achieve effective communication. In this way, by predicting the communication quality in advance and scheduling it in advance, the user's communication experience can be guaranteed.
[0146] Specific limitations regarding the communication device can be found in the limitations regarding the communication method above, and will not be repeated here. Each module in the aforementioned communication device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.
[0147] In addition, this application also provides an electronic device, such as Figure 11 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically: The electronic device may include components such as a processor 1101 with one or more processing cores and a memory 1102 of one or more computer-readable storage media. Those skilled in the art will understand that... Figure 11 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 1101 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1102, and by calling data stored in the memory 1102, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 1101 may include one or more processing cores; preferably, the processor 1101 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1101.
[0148] The memory 1102 can be used to store software programs and modules. The processor 1101 executes various functional applications and data processing by running the software programs and modules stored in the memory 1102. The memory 1102 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1102 may also include a memory controller to provide the processor 1101 with access to the memory 1102.
[0149] In one embodiment, the electronic device further includes a power supply 1103 for supplying power to the various components. Preferably, the power supply 1103 can be logically connected to the processor 1101 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 1103 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0150] In one embodiment, the electronic device may further include an input unit 1104, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0151] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 1101 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 1102 according to the following instructions, and the processor 1101 runs the applications stored in the memory 1102, thereby realizing the steps in any of the communication methods provided in the embodiments of this application.
[0152] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0153] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described in any embodiment of this application.
[0154] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of this application.
[0155] In some embodiments, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the methods described in any embodiment of this application.
[0156] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0157] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0158] Therefore, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the communication methods provided in this application.
[0159] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0160] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0161] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the communication methods provided in this application, the beneficial effects that any of the communication methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0162] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0163] The foregoing has provided a detailed description of a communication method, system, device, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A communication method, characterized in that, Applied to the sending end; the method includes: Predict future Quality of Service (QoS) vectors based on current communication characteristics; Based on the future QoS vector, the collected audio and / or video streams are processed to obtain the target data; The target data is sent to the receiving end so that the receiving end can input the target data into the generation model to obtain the communication content.
2. The method according to claim 1, characterized in that, The current communication characteristics include: network characteristics, media characteristics, and / or environmental characteristics; Based on current communication characteristics, predict future QoS vectors, including: The current communication features are input into the temporal convolutional network and long short-term memory network of the dual-window prediction model to obtain sudden change features and long-term trend features. The sudden change features and the long-term trend features are fused based on a cross-modal attention mechanism to obtain fused features; Based on the fused features, the future QoS vector is predicted.
3. The method according to claim 1, characterized in that, The step of processing the acquired audio and / or video streams based on future QoS vectors to obtain target data includes: Determine the future communication mode based on the future QoS vector; According to the future communication mode, the acquired audio stream and / or video stream are processed to obtain the target data.
4. The method according to claim 3, characterized in that, Determining the future communication mode based on the future QoS vector includes: Under the future QoS vector, obtain the audio loss value, video loss value and generation delay corresponding to different communication modes; Based on the audio loss value, the video loss value, and the generation delay corresponding to each of the different communication modes, determine the QoS loss value corresponding to each of the different communication modes under the future QoS vector; The communication mode with the smallest QoS loss value among different communication modes is determined as the future communication mode.
5. The method according to claim 3, characterized in that, Determining the future communication mode based on the future QoS vector includes: Obtain the current QoS vector and the hysteresis interval corresponding to the current communication mode; The future communication mode is determined when the future QoS vector is greater than the upper limit of the hysteresis interval or less than the lower limit of the hysteresis interval.
6. The method according to claim 3, characterized in that, The future communication modes include at least: audio degradation mode, video degradation mode and / or background degradation mode; The processing of the acquired audio stream and / or video stream according to the future communication mode includes one or more of the following steps: In the audio degradation mode, the acquired audio stream is subjected to compression processing, voiceprint feature extraction processing, speech-to-text processing, and / or semantic summary text generation processing. In the video degradation mode, the acquired video stream is subjected to resolution reduction processing, extraction of avatar parameters at multiple times, and / or extraction of static avatar parameters. In the background degradation mode, a background image is extracted from the acquired video stream, and a background label is generated based on the background image.
7. The method according to claim 1, characterized in that, Before sending the target data to the receiving end, the method further includes: Obtain the session identifier and time slice information corresponding to the target data, and convert the session identifier and time slice information into digital watermarks of multiple modalities; The digital watermarks of multiple modalities are respectively added to the target data with the same modality as the digital watermark, so as to obtain the target data with added digital watermarks, so that the receiving end can perform cross-modal verification based on the digital watermarks extracted from the target data.
8. The method according to claim 1, characterized in that, The sending end is a participant in a multi-terminal conference; The step of processing the acquired audio and / or video streams based on future QoS vectors to obtain target data includes: Obtain its own first priority, and determine a first future communication mode based on the future QoS vector and the first priority; According to the first future communication mode, the collected audio stream and / or video stream are processed to obtain the target data.
9. A communication method, characterized in that, Applied to the receiving end; the method includes: Receive the target data sent by the sending end; The target data is input into the generation model to obtain the communication content; The target data is obtained by the sending end predicting the future QoS vector based on the current communication characteristics, and processing the collected audio stream and / or video stream according to the future QoS vector.
10. The method according to claim 9, characterized in that, The generative model includes: an audio generation network, an avatar generation network, and a background generation network; the target data includes: speech text, voiceprint features, avatar parameters at multiple time points, static avatar parameters, and / or background labels; The step of inputting the target data into the generation model to obtain the communication content includes one or more of the following steps: The audio generation network is input with the spoken text and the voiceprint features to predict the audio prosody, and then generates reconstructed audio based on the audio prosody, the spoken text, and the voiceprint features. The avatar parameters at multiple times are input into the avatar generation network to generate dynamic avatars; The static avatar parameters are input into the avatar generation network to generate a static avatar. The background label is input into the background generation network to generate a background image.
11. The method according to claim 9, characterized in that, The generative model includes: an audio generation network and an avatar generation network; the training steps of the generative model include: Obtain the audio features output by the audio generation network to be trained; Obtain the lip features corresponding to the audio features output by the head-image generation network to be trained; Based on the differences between the audio features and the lip-sync features, a speech-lip-sync consistency loss function is established; The generative model to be trained is trained according to the speech-lip consistency loss function to obtain the trained generative model.
12. The method according to claim 9, characterized in that, The step of inputting the target data into the generation model to obtain the communication content includes: After receiving the target data with added digital watermark, extract the digital watermarks of the target data in multiple modalities; Cross-modal verification is performed based on the digital watermark of the target data; When the cross-modal validation passes, the target data is input into the generative model to obtain the communication content; The digital watermark in the target data is obtained by the sending end from converting multiple modal digital watermarks based on the session identifier and time slice information corresponding to the target data, and then adding the multiple modal digital watermarks to the target data with the same modality as the digital watermark.
13. The method according to claim 9, characterized in that, The receiving end is a participant in a multi-terminal conference; The step of inputting the target data into the generation model to obtain the communication content includes: Predict its own future QoS vector and obtain its own second priority; Based on its own future QoS vector and the second priority, a second future communication mode is determined; According to the second future communication mode, the received target data is selected; The selected target data is input into the generation model to obtain the communication content.
14. A communication system, characterized in that, include: The sending end is used to predict the future QoS vector based on the current communication characteristics; process the collected audio stream and / or video stream according to the future QoS vector to obtain target data, and send the target data to the receiving end; The receiving end is used to input the target data into the generation model to obtain the communication content.
15. The system according to claim 14, characterized in that, The current communication characteristics include: network characteristics, media characteristics, and / or environmental characteristics; The sending end predicts the future QoS vector based on current communication characteristics, including: The transmitting end inputs the current communication features into the temporal convolutional network and long short-term memory network of the dual-window prediction model to obtain sudden change features and long-term trend features. The sudden change features and the long-term trend features are fused based on a cross-modal attention mechanism to obtain fused features; Based on the fused features, the future QoS vector is predicted.
16. The system according to claim 14, characterized in that, The transmitting end processes the acquired audio and / or video streams according to the future QoS vector to obtain target data, including: The sending end determines the future communication mode based on the future QoS vector; The transmitting end processes the acquired audio stream and / or video stream according to the future communication mode to obtain the target data.
17. The system according to claim 16, characterized in that, The sending end determines the future communication mode based on the future QoS vector, including: The transmitting end obtains the audio loss value, video loss value, and generation delay corresponding to different communication modes under the future QoS vector; Based on the audio loss value, the video loss value, and the generation delay corresponding to each of the different communication modes, determine the QoS loss value corresponding to each of the different communication modes under the future QoS vector; The communication mode with the smallest QoS loss value among different communication modes is determined as the future communication mode.
18. The system according to claim 16, characterized in that, The sending end determines the future communication mode based on the future QoS vector, including: The sending end obtains the current QoS vector and the hysteresis interval corresponding to the current communication mode; The future communication mode is determined when the future QoS vector is greater than the upper limit of the hysteresis interval or less than the lower limit of the hysteresis interval.
19. The system according to claim 16, characterized in that, The future communication modes include at least: audio degradation mode, video degradation mode and / or background degradation mode; The transmitting end processes the acquired audio stream and / or video stream according to the future communication mode, including one or more of the following steps: In the audio degradation mode, the acquired audio stream is subjected to compression processing, voiceprint feature extraction processing, speech-to-text processing, and / or semantic summary text generation processing. In the video degradation mode, the acquired video stream is subjected to resolution reduction processing, extraction of avatar parameters at multiple times, and / or extraction of static avatar parameters. In the background degradation mode, a background image is extracted from the acquired video stream, and a background label is generated based on the background image.
20. The system according to claim 14, characterized in that, The generative model includes: an audio generation network, an avatar generation network, and a background generation network; the target data includes: speech text, voiceprint features, avatar parameters at multiple time points, static avatar parameters, and / or background labels; The receiving end inputs the target data into the generation model to obtain the communication content, including one or more of the following steps: The receiving end inputs the speech text and the voiceprint features into the audio generation network to predict the audio prosody, and generates reconstructed audio based on the audio prosody, the text, and the voiceprint features; The avatar parameters at multiple times are input into the avatar generation network to generate dynamic avatars; The static avatar parameters are input into the avatar generation network to generate a static avatar. The background label is input into the background generation network to generate a background image.
21. The system according to claim 14, characterized in that, The generative model includes: an audio generation network and an avatar generation network; the training steps of the generative model include: Obtain the audio features output by the audio generation network to be trained; Obtain the lip features corresponding to the audio features output by the head-image generation network to be trained; Based on the differences between the audio features and the lip-sync features, a speech-lip-sync consistency loss function is established; The generative model to be trained is trained according to the speech-lip consistency loss function to obtain the trained generative model.
22. The system according to claim 14, characterized in that, The sending end is used to obtain the session identifier and time slice information corresponding to the target data, and convert the session identifier and time slice information into a digital watermark with multiple modalities. The digital watermarks of multiple modalities are respectively added to the target data with the same modality as the digital watermarks to obtain the target data with added digital watermarks; The receiving end inputs the target data into the generation model to obtain the communication content, including: After receiving the target data with added digital watermarks corresponding to the same time slice, the receiving end extracts the digital watermarks of the target data in multiple modalities. Cross-modal verification is performed based on the digital watermark of the target data; When the cross-modal validation passes, the target data is input into the generative model to obtain the communication content.
23. The system according to claim 14, characterized in that, The sending end and the receiving end are participants in a multi-terminal conference; The transmitting end processes the acquired audio and / or video streams according to the future QoS vector to obtain target data, including: The sending end obtains its own first priority and determines a first future communication mode based on the future QoS vector and the first priority; The transmitting end processes the acquired audio stream and / or video stream according to the first future communication mode to obtain the target data; The receiving end inputs the target data into the generation model to obtain the communication content, including: The receiving end predicts its own future QoS vector and obtains its own second priority; The receiving end determines the second future communication mode based on its own future QoS vector and the second priority; The receiving end selects the received target data according to the second future communication mode, and inputs the selected target data into the generation model to obtain the communication content.
24. A communication device, characterized in that, Applied to the transmitting end; the device includes: The prediction module is used to predict future QoS vectors based on current communication characteristics; The processing module is used to process the acquired audio and / or video streams according to the future QoS vector to obtain the target data; The sending module is used to send the target data to the receiving end, so that the receiving end can input the target data into the generation model to obtain the communication content.
25. A communication device, characterized in that, Applied to the receiving end; the device includes: The receiving module is used to receive the target data sent by the sending end; The input module is used to input the target data into the generation model to obtain the communication content; The target data is obtained by the sending end predicting the future QoS vector based on the current communication characteristics, and processing the collected audio stream and / or video stream according to the future QoS vector.
26. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the communication method as described in any one of claims 1 to 13.
27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the communication method as described in any one of claims 1 to 13.