Video conference quality optimization method and system in weak network scene, medium and terminal
By acquiring and converting speaker data into small amounts of information in a weak network environment, and restoring video data using voiceprints and portrait models, the quality issues of video conferencing in weak network environments are resolved, improving user experience and data transmission efficiency.
Patent Information
- Application Number
- CN202510726764.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
Smart Images

Figure CN120676115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio and video transmission in weak network scenarios, and in particular to a method, system, medium and terminal for optimizing video conferencing quality in weak network scenarios. Background Art
[0002] With the rapid advancement of internet technology, video conferencing has become widely used in many fields. However, in practice, video conferencing is susceptible to unstable network conditions due to factors such as relatively underdeveloped network infrastructure in some areas, the large number of devices connected simultaneously, large amounts of video data, and insufficient terminal device performance. Data transmission delays, freezes, and even interruptions are common, especially in environments with poor network conditions, significantly reducing the quality and effectiveness of video conferencing.
[0003] Currently, a common strategy is to reduce data volume by reducing video resolution and frame rate. However, such solutions inevitably lead to a significantly reduced conference experience on the receiving end and fail to meet the needs of users in weak network environments. Summary of the Invention
[0004] The present invention provides a method, system, medium and terminal for optimizing video conferencing quality in a weak network scenario. In a weak network scenario, it not only ensures that the user experience of the receiving end is not damaged, but also solves video conferencing quality problems such as lag and packet loss caused by excessive data transmission.
[0005] In a first aspect, an embodiment of the present invention provides a method for optimizing video conferencing quality in a weak network scenario, comprising:
[0006] During a video conference, first data information corresponding to a speaker's human behavior characteristics and the speaker's voice data are obtained in real time, and whether the current system is in a weak network scenario is determined in real time; wherein the first data information includes speaker identity information and speaker behavior characteristic information;
[0007] If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information and the second data information, and combines the speaker voice and the speaker behavior to restore and generate the audio and video data corresponding to the video conference.
[0008] In an embodiment of the present invention, during a video conference, first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data are obtained in real time. The first data information includes the speaker's identity information and the speaker's behavior characteristic information, and a determination is made in real time as to whether the current system is in a weak network scenario. If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to a receiving end. The first data information and the second data information include text, code, or other structured symbolic information (such as a binary vector, a feature parameter sequence, etc.). Since these data information are much smaller than the amount of video and voice data, in a weak network scenario, the video conference voice data and other relevant data such as the speaker's human behavior characteristics are converted into data information, and only the first data information corresponding to the speaker's human behavior characteristics and the second data information corresponding to the speaker's voice data are transmitted. Compared to sending a large amount of data such as the video conference voice data and other relevant data to the receiving end in a weak network scenario, this can greatly reduce the amount of data transmitted between the sending end and the receiving end during the video conference, reduce the system network burden, avoid problems such as lag and packet loss caused by large amounts of data, and thereby improve the user experience of users participating in the video conference. Furthermore, when the receiving end receives the first data information and the second data information, it calls the corresponding speaker voiceprint information based on the speaker identity information in the first data information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information in combination with the first data information and the second data information, so as to more realistically restore the unique voice characteristics and the speaker's posture when speaking displayed in the meeting, and combine the speaker's voice and speaker behavior to restore and generate the video data corresponding to the video conference, so as to quickly and accurately convey the video data corresponding to the video conference to the receiving end, so that the receiving end can more accurately understand the speaker's emotions, tone and intentions, so as to improve the conference experience of the receiving end users while solving the problems faced by data transmission in weak network environments.
[0009] As a preferred example of the first aspect, if the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information, and the second data information, specifically:
[0010] If the current system is in a weak network scenario, converting the voice data into corresponding second data information, and transmitting the first data information and the second data information to a receiving end, so that the receiving end retrieves the speaker voiceprint information corresponding to the speaker identity information from a database based on the speaker identity information;
[0011] The receiving end generates corresponding speaker voice according to the speaker voiceprint information and the second data information by using natural language processing technology and a speech synthesis engine.
[0012] In this preferred example, personalized conference restoration is achieved by calling the speaker's voiceprint information at the receiving end, and the receiving end uses natural language processing technology and a speech synthesis engine to generate the corresponding speaker's voice.
[0013] As a preferred example of the first aspect, if the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information, and the second data information, specifically:
[0014] If the current system is in a weak network scenario, converting the voice data into corresponding second data information, and transmitting the first data information and the second data information to a receiving end, so that the receiving end calls a speaker portrait model corresponding to the speaker identity information from a database based on the speaker identity information;
[0015] The receiving end generates a continuous action of the corresponding portrait model according to the portrait model and the speaker behavior feature information, and the speaker behavior is constituted based on the continuous action of the portrait model.
[0016] In this preferred example, by calling the speaker's portrait model at the receiving end, and generating the corresponding continuous movements of the portrait model based on the portrait model and the speaker's behavioral feature information, and based on the continuous movements of the portrait model, the speaker's behavior is constructed, thereby generating unique voice and behavioral performance for each speaker and realizing personalized conference restoration.
[0017] As a preferred example of the first aspect, the receiving end generates, based on the portrait model and the speaker behavior feature information, a continuous action of the corresponding portrait model, and constitutes the speaker behavior based on the continuous action of the portrait model, specifically:
[0018] Based on the speaker's behavior feature information, obtaining a plurality of human behavior guiding words of the speaker and the order of all the human behavior guiding words; wherein the human behavior guiding words include expression guiding words and / or gesture guiding words;
[0019] Based on a plurality of human behavior guiding words of the speaker, calling from a database the limb change sequences corresponding to the human behavior guiding words; wherein the limb change sequences are pre-generated using motion capture and animation generation technology and stored in the database;
[0020] According to the order of all the human behavior guiding words, the limb change sequence corresponding to each of the human behavior guiding words is demonstrated on the portrait model in turn to form the continuous movement of the portrait model, and based on the continuous movement of the portrait model, the speaker behavior is constructed.
[0021] In this preferred example, by obtaining human behavior guide words (such as expression guide words and gesture guide words) and their order, the corresponding limb change sequence is called from the local database, and the limb change sequence is demonstrated in sequence according to the order of the human behavior guide words, ensuring that the actions presented on the portrait model are continuous and consistent with the speaker's behavioral logic in the actual meeting. The various subtle behaviors of the speaker can be accurately restored on the portrait model, so that the restored speaker behavior looks natural and smooth, thereby enhancing the realism and immersion of the video conference.
[0022] As a preferred example of the first aspect, if the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, specifically:
[0023] If the current system is in a weak network scenario, corresponding voice data is obtained based on the voice signal, and then a pause and silence duration of the voice data is detected. Based on the detection result, the voice data is preprocessed to obtain a corresponding vector sequence; wherein the voice signal includes the voice data, and the vector sequence includes at least one voice vector;
[0024] By calling a speech recognition model, speech recognition is performed on the speech vectors in the vector sequence, and second data information corresponding to the speech data is generated by combining the detection result and the speech recognition result.
[0025] In this preferred example, by detecting the pause and silence duration of the speech data and incorporating it into subsequent processing, different semantic contents are avoided from being mixed together for recognition, which helps to more accurately divide sentence boundaries, thereby improving the accuracy of speech recognition.
[0026] As a preferred example of the first aspect, in the process of the video conference, the first data information corresponding to the human behavior characteristics of the speaker and the voice data of the speaker are obtained in real time, specifically:
[0027] During the video conference, human behavior characteristics corresponding to the speaker and the speaker's voice data are obtained in real time; wherein the human behavior characteristics include facial features and / or body movements;
[0028] Comparing the facial features with pre-stored user data to obtain corresponding speaker identity information;
[0029] Combining the facial features and the voice data, the speaker's emotional features are extracted in real time to obtain corresponding speaker emotional information, and key frame extraction and pattern matching are performed on the body movements to obtain corresponding special gesture signals. Then, the speaker's emotional information and the special gesture signals are combined to constitute the first data information corresponding to the speaker's human behavior characteristics.
[0030] In this preferred example, the speaker's tone and facial features can convey the speaker's emotions, and emotions are an important part of communication. Different emotions may affect the understanding of the same words, and gestures often play a role in assisting expression in communication. Therefore, combined with facial features and voice data, the speaker's emotional features are extracted in real time to obtain the corresponding speaker's emotional information, and key frame extraction and pattern matching are performed on the body movements to obtain the corresponding special gesture signals. Then, the speaker's emotional information and special gesture signals are combined to form the first data information corresponding to the speaker's human behavior characteristics, so that the receiving end can accurately restore the video data corresponding to the video conference based on the speaker's emotional information and special gesture signals in the first data information, so that the participants at the receiving end can accurately understand the speaker's expression intentions, thereby improving the efficiency and quality of conference communication.
[0031] As a preferred example of the first aspect, the real-time determination of whether the current system is in a weak network scenario is specifically as follows:
[0032] During the video conference, network indicator values corresponding to several system network indicators of the current system are obtained in real time, and each of the network indicator values is compared with a preset limit range of the corresponding system network indicator;
[0033] If any of the network indicator values does not meet the preset limit range of the corresponding system network indicator, it is determined that the current system is in a weak network scenario;
[0034] If all the network indicator values are within the preset limit range of the corresponding system network indicator, it is determined that the current system is not in a weak network scenario.
[0035] In this preferred example, the network status changes dynamically. By obtaining the network indicator values corresponding to several system network indicators of the current system in real time during the video conference, and comparing each network indicator value with the preset limit range of the corresponding system network indicator, network anomalies can be quickly discovered, and corresponding measures can be taken to deal with the current weak network scenario to ensure that the video conference can be carried out continuously and stably.
[0036] In a second aspect, an embodiment of the present application further provides a system for optimizing video conferencing quality in a weak network scenario, including:
[0037] An acquisition and judgment module is used to obtain, in real time during the video conference, first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data, and to determine in real time whether the current system is in a weak network scenario; wherein the first data information includes the speaker's identity information and the speaker's behavior characteristic information;
[0038] A video restoration module is used to convert the voice data into corresponding second data information if the current system is in a weak network scenario, and transmit the first data information and the second data information to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information and the second data information, and combines the speaker voice and the speaker behavior to restore and generate the audio and video data corresponding to the video conference.
[0039] As a preferred example of the second aspect, the video restoration module specifically includes:
[0040] an information transmission unit, configured to, if the current system is in a weak network scenario, convert the voice data into corresponding second data information, and transmit the first data information and the second data information to a receiving end, so that the receiving end retrieves the speaker voiceprint information and portrait model corresponding to the speaker identity information from a database based on the speaker identity information;
[0041] The video restoration unit is used to generate the corresponding speaker voice based on the speaker voiceprint information and the second data information through the receiving end using natural language processing technology and a speech synthesis engine; generate the corresponding continuous movement of the portrait model based on the portrait model and the speaker behavior feature information through the receiving end, and constitute the speaker behavior based on the continuous movement of the portrait model; and restore and generate the audio and video data corresponding to the video conference in combination with the speaker voice and the speaker behavior.
[0042] In order to solve the same technical problem, the present invention also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for optimizing video conferencing quality in a weak network scenario.
[0043] In order to solve the same technical problem, the present invention also provides a terminal, including a processor, a memory and a computer program stored in the memory; wherein, the computer program can be executed by the processor to implement the method for optimizing video conferencing quality in a weak network scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 This is a flowchart of a method for optimizing video conferencing quality in a weak network scenario provided by some embodiments of the present invention;
[0046] Figure 2 This is a structural diagram of a device for optimizing video conferencing quality in a weak network scenario provided by some embodiments of the present invention. DETAILED DESCRIPTION
[0047] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.
[0049] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.
[0050] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0051] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0052] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0053] In the description of the embodiments of the present application, unless otherwise expressly specified or limited, technical terms such as "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; internal connections between two components or interactions between two components. Those skilled in the art can understand the specific meanings of the above terms in the embodiments of the present application based on specific circumstances.
[0054] With the rapid advancement of internet technology, video conferencing has become widely used in many fields. However, in actual video conferencing, data transmission delays, freezes, and even interruptions often occur due to unstable network conditions. This is particularly prominent in environments with poor network conditions, significantly reducing the quality and effectiveness of video conferencing.
[0055] Currently, a common strategy is to reduce data volume by reducing video resolution and frame rate. However, such solutions inevitably compromise the conferencing experience on the receiving end and fail to meet user needs in weak network environments. Therefore, a new technical solution is urgently needed to address extremely weak network scenarios and improve the efficiency and quality of voice information transmission. This solution must ensure that the receiving end's user experience is not compromised while also properly addressing the challenges of data transmission in weak network environments.
[0056] See also Figure 1 To address the difficulties of existing technologies in weak network scenarios, ensuring that the user experience of the receiving end is not compromised while also solving video conferencing quality issues such as lag and packet loss caused by excessive data transmission, an embodiment of the present invention provides a method for optimizing video conferencing quality in weak network scenarios, including:
[0057] S1. During a video conference, obtaining, in real time, first data information corresponding to a speaker's human behavior characteristics and the speaker's voice data, and determining in real time whether the current system is in a weak network scenario; wherein the first data information includes speaker identity information and speaker behavior characteristic information;
[0058] S2. If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information and the second data information, and combines the speaker voice and the speaker behavior to restore and generate the audio and video data corresponding to the video conference.
[0059] Specifically, the speaker's behavior can be reflected as a continuous action sequence, and also includes generating a labeled output representing the behavior based on the received first data information. The labeled output includes but is not limited to barrage text, emoticons, short animation pop-ups, and other forms.
[0060] Furthermore, in some embodiments of the present application, if the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information, and the second data information, specifically:
[0061] If the current system is in a weak network scenario, converting the voice data into corresponding second data information, and transmitting the first data information and the second data information to a receiving end, so that the receiving end retrieves the speaker voiceprint information corresponding to the speaker identity information from a database based on the speaker identity information;
[0062] The receiving end generates corresponding speaker voice according to the speaker voiceprint information and the second data information by using natural language processing technology and a speech synthesis engine.
[0063] Specifically, in order to fully explain the above steps, the following scheme can be used as an example for illustration:
[0064] Using the speaker's identity information, the speaker's voiceprint, generated using deep learning, is retrieved from a local database. Natural language processing technology and a speech synthesis engine are then used to convert the second data into speech for playback. The specific process is as follows: First, the voice information of all possible participants is recorded in advance. The voice information is segmented and features extracted, and converted into voiceprint feature vectors. A database is then established to record the participants and their voiceprint information. Next, the receiving end converts the corresponding speaker's voice into data, such as text, codes, and symbols. This data is processed based on the receiving end's information. For example, if the receiving end's nationality or language differs from the speaker's, the received text is translated. Finally, using the speaker's pre-recorded voiceprint information, or using a default voice if the speaker hasn't recorded in advance, the processed text is synthesized into speech (TTS). This speech synthesis uses a trained TTS model, which takes as input a text and voiceprint features and outputs a speech waveform that approximates the input voiceprint features. The speech parameters are user-configurable.
[0065] In this way, personalized conference restoration is achieved by calling the speaker's voiceprint information at the receiving end, and the receiving end uses natural language processing technology and speech synthesis engine to generate the corresponding speaker's voice.
[0066] Furthermore, in some embodiments of the present application, if the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information, and the second data information, specifically:
[0067] If the current system is in a weak network scenario, converting the voice data into corresponding second data information, and transmitting the first data information and the second data information to a receiving end, so that the receiving end calls a speaker portrait model corresponding to the speaker identity information from a database based on the speaker identity information;
[0068] The receiving end generates a continuous action of the corresponding portrait model according to the portrait model and the speaker behavior feature information, and the speaker behavior is constituted based on the continuous action of the portrait model.
[0069] Specifically, in order to fully explain the above steps, the following scheme can be used as an example:
[0070] Using the speaker's identity information, a portrait model generated using 3D modeling technology is retrieved from a local database. Based on the descriptors from the data analysis results, expression or gesture guides are obtained. Based on these expression or gesture guides, motion capture and animation generation technologies are used to achieve the digital human's natural and smooth target movements. The specific process is as follows: First, a preset body movement sequence is generated using motion capture or animation generation technology: for example, an animation of a laughing expression, an animation of raising a hand, or an animation of a prolonged speech. After receiving the expression, gesture, and body movement information, the receiver uses a pre-programmed judgment module to determine which body movement sequence should be used and then displays it on the 3D model. Furthermore, for switching between different movements, the matrix of the last frame of the body animation for the current movement is interpolated with the matrix of the first frame of the next movement to ensure smooth transitions.
[0071] In this way, by calling the speaker's portrait model at the receiving end, and generating the corresponding continuous movements of the portrait model based on the portrait model and the speaker's behavioral feature information, and based on the continuous movements of the portrait model, the speaker's behavior is constructed, thereby generating unique voice and behavioral performance for each speaker and realizing personalized conference restoration.
[0072] Furthermore, in some embodiments of the present application, the receiving end generates a continuous action of the corresponding portrait model based on the portrait model and the speaker behavior feature information, and constitutes the speaker behavior based on the continuous action of the portrait model, specifically:
[0073] Based on the speaker's behavior feature information, obtaining a plurality of human behavior guiding words of the speaker and the order of all the human behavior guiding words; wherein the human behavior guiding words include expression guiding words and / or gesture guiding words;
[0074] Based on a plurality of human behavior guiding words of the speaker, calling from a database the limb change sequences corresponding to the human behavior guiding words; wherein the limb change sequences are pre-generated using motion capture and animation generation technology and stored in the database;
[0075] According to the order of all the human behavior guiding words, the limb change sequence corresponding to each of the human behavior guiding words is demonstrated on the portrait model in turn to form the continuous movement of the portrait model, and based on the continuous movement of the portrait model, the speaker behavior is constructed.
[0076] In this way, by obtaining human behavior guide words (such as expression guide words and / or gesture guide words) and their order, the corresponding limb change sequence is called from the local database, and the limb change sequence is demonstrated in sequence according to the order of the human behavior guide words, ensuring that the actions presented on the portrait model are continuous and consistent with the speaker's behavioral logic in the actual meeting. The various subtle behaviors of the speaker can be accurately restored on the portrait model, making the restored speaker behavior look natural and smooth, thereby enhancing the realism and immersion of the video conference.
[0077] Furthermore, in some embodiments of the present application, if the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, specifically:
[0078] If the current system is in a weak network scenario, corresponding voice data is obtained based on the voice signal, and then a pause and silence duration of the voice data is detected. Based on the detection result, the voice data is preprocessed to obtain a corresponding vector sequence; wherein the voice signal includes the voice data, and the vector sequence includes at least one voice vector;
[0079] By calling a speech recognition model, speech recognition is performed on the speech vectors in the vector sequence, and second data information corresponding to the speech data is generated by combining the detection result and the speech recognition result.
[0080] In this way, by detecting the pause and silence duration of the speech data and incorporating it into subsequent processing, we can avoid confusing different semantic contents together for recognition, help to more accurately divide sentence boundaries, and thus improve the accuracy of speech recognition.
[0081] The conversion of the voice data into the corresponding second data information is preferably performed locally by the sending end. When the bandwidth conditions of the sending end permit, the voice data can also be optionally sent to the server for the second data information conversion. In some embodiments of the present application, when the conference scenario requires end-to-end encryption (such as financial and government meetings), the voice data conversion is performed locally by the sending end, which can avoid the risk of data leakage caused by data transmission through the server, while ensuring the security of sensitive data such as voiceprint information and user behavior characteristics.
[0082] Furthermore, in some embodiments of the present application, during the video conference, first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data are obtained in real time, specifically: the first data information includes the speaker's identity information and the speaker's behavior characteristic information; the human behavior characteristics include facial features and / or body movements; the speaker's identity information is obtained through technical means such as biometric recognition (such as facial features, voiceprints) or participant identification (such as participant ID).
[0083] Combining the facial features and the voice data, the speaker's emotional features are extracted in real time to obtain corresponding speaker emotional information, and key frame extraction and pattern matching are performed on the body movements to obtain corresponding special gesture signals. Then, the speaker's emotional information and the special gesture signals are combined to constitute the first data information corresponding to the speaker's human behavior characteristics.
[0084] Specifically, the real-time emotional feature extraction of the speaker to obtain the corresponding emotional information of the speaker can be illustrated by taking the following scheme as an example:
[0085] Take the extraction of emotional features as an example. For datasets, emotion recognition is a classification problem in deep learning. Commonly used datasets categorize seven main emotions: anger, disgust, fear, happiness, sadness, surprise, and neutral. Data labels can also be processed differently for different scenarios. For example, in scenarios where classification doesn't require this level of accuracy, emotions can be categorized as positive, negative, and neutral. For the structure of the emotion model, lightweight networks with a small number of parameters are preferred as the backbone network for image feature extraction, such as the MobileNet series, Inception series, ResNet series, and ShuffleNet series. After feature extraction, a fully connected layer or global average pooling layer is added to map the extracted backbone features to the classification labels. For model input, facial image data extracted by the face detection model is used as the input for the emotion recognition model.
[0086] In this way, the speaker's tone and facial features can convey the speaker's emotions, and emotions are an important part of communication. Different emotions may affect the understanding of the same words. Gestures often play a role in assisting expression in communication. Therefore, combined with facial features and voice data, the speaker's emotional features are extracted in real time to obtain the corresponding speaker's emotional information, and key frame extraction and pattern matching are performed on the body movements to obtain the corresponding special gesture signals. Then, the speaker's emotional information and special gesture signals are combined to form the first data information corresponding to the speaker's human behavior characteristics, so that the receiving end can accurately restore the video data corresponding to the video conference based on the speaker's emotional information and special gesture signals in the first data information, so that the participants at the receiving end can accurately understand the speaker's expression intentions, thereby improving the efficiency and quality of conference communication.
[0087] Furthermore, in some embodiments of the present application, the real-time determination of whether the current system is in a weak network scenario is specifically as follows:
[0088] During the video conference, network indicator values corresponding to several system network indicators of the current system are obtained in real time, and each of the network indicator values is compared with a preset limit range of the corresponding system network indicator;
[0089] If any of the network indicator values does not meet the preset limit range of the corresponding system network indicator, it is determined that the current system is in a weak network scenario;
[0090] If all the network indicator values are within the preset limit range of the corresponding system network indicator, it is determined that the current system is not in a weak network scenario.
[0091] Specifically, in order to fully explain the above judgment process, the following scheme is used as an example for illustration:
[0092] When the user experience is clearly poor, we monitor network jitter, bandwidth, packet loss, and other metrics. If we find that network jitter exceeds 120ms or packet loss exceeds 50%, we can use 120ms of jitter or 50% of packet loss as the standard for a weak network in this scenario. Later, during the video conference, we monitor the receiving end's network bandwidth, latency, and packet loss rate in real time. If the network bandwidth falls below a preset threshold, the latency exceeds a certain limit, or the packet loss rate exceeds a specific percentage, we can determine the current weak network status based on the network metrics at both ends.
[0093] Since network conditions change dynamically, by obtaining the network indicator values corresponding to several system network indicators of the current system in real time during the video conference, and comparing each network indicator value with the preset limit range of the corresponding system network indicator, network anomalies can be quickly discovered, and corresponding measures can be taken to deal with the current weak network scenario to ensure that the video conference can be carried out continuously and stably.
[0094] In summary, by implementing an embodiment of the present invention, during a video conference, first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data are obtained in real time, the first data information including the speaker's identity information and the speaker's behavior characteristic information, and it is determined in real time whether the current system is in a weak network scenario. If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end. Since the amount of data information is much smaller than that of video and voice data, in a weak network scenario, the voice data of the video conference and other relevant data such as the speaker's human behavior characteristics are converted into data information, and only the first data information corresponding to the speaker's human behavior characteristics and the second data information corresponding to the speaker's voice data are transmitted. Compared with sending a large amount of data such as the voice data and other relevant data of the video conference to the receiving end in a weak network scenario, the amount of data transmission between the sending end and the receiving end during the video conference can be greatly reduced, and the system network burden can be reduced to avoid problems such as freezes and packet loss caused by large amounts of data, thereby improving the user experience of users participating in the video conference. Furthermore, when the receiving end receives the first data information and the second data information, it calls the corresponding speaker voiceprint information based on the speaker identity information in the first data information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information in combination with the first data information and the second data information, so as to more realistically restore the unique voice characteristics and the speaker's posture when speaking displayed in the meeting, and combine the speaker's voice and speaker behavior to restore and generate the video data corresponding to the video conference, so as to quickly and accurately convey the video data corresponding to the video conference to the receiving end, so that the receiving end can more accurately understand the speaker's emotions, tone and intentions, so as to improve the conference experience of the receiving end users while solving the problems faced by data transmission in weak network environments.
[0095] Example 2
[0096] like Figure 2 As shown, based on the above method embodiment, a corresponding system embodiment is provided;
[0097] An embodiment of the present invention provides a system for optimizing video conferencing quality in a weak network scenario, comprising: an acquisition and judgment module 11 and a video restoration module;
[0098] Furthermore, in some embodiments of the present application, the acquisition and judgment module 11 is used to obtain first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data in real time during the video conference, and to determine in real time whether the current system is in a weak network scenario; wherein, the first data information includes speaker identity information and speaker behavior characteristic information; the video restoration module 12 is used to convert the voice data into corresponding second data information if the current system is in a weak network scenario, and transmit the first data information and the second data information to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information and the second data information, and restores and generates the audio and video data corresponding to the video conference in combination with the speaker voice and the speaker behavior.
[0099] Furthermore, in some embodiments of the present application, the video restoration module 12 includes an information transmission unit and a video restoration unit; the information transmission unit is used to convert the voice data into corresponding second data information if the current system is in a weak network scenario, and transmit the first data information and the second data information to the receiving end, so that the receiving end calls the speaker voiceprint information and portrait model corresponding to the speaker identity information from the database based on the speaker identity information; the video restoration unit is used to generate the corresponding speaker voice according to the speaker voiceprint information and the second data information through the receiving end using natural language processing technology and a speech synthesis engine; through the receiving end, generate a labeled output of the corresponding behavioral feature according to the portrait model and the speaker behavior feature information; preferably, the continuous action of the portrait model, and based on the continuous action of the portrait model, the speaker behavior is constituted; combined with the speaker voice and the speaker behavior, restore and generate the audio and video data corresponding to the video conference.
[0100] In summary, by implementing an embodiment of the present invention, during a video conference, first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data are obtained in real time, the first data information including the speaker's identity information and the speaker's behavior characteristic information, and it is determined in real time whether the current system is in a weak network scenario. If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end. Since the amount of data information is much smaller than that of video and voice data, in a weak network scenario, the voice data of the video conference and other relevant data such as the speaker's human behavior characteristics are converted into data information, and only the first data information corresponding to the speaker's human behavior characteristics and the second data information corresponding to the speaker's voice data are transmitted. Compared with sending a large amount of data such as the voice data and other relevant data of the video conference to the receiving end in a weak network scenario, the amount of data transmission between the sending end and the receiving end during the video conference can be greatly reduced, and the system network burden can be reduced to avoid problems such as freezes and packet loss caused by large amounts of data, thereby improving the user experience of users participating in the video conference. Furthermore, when the receiving end receives the first data information and the second data information, it calls the corresponding speaker voiceprint information based on the speaker identity information in the first data information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information in combination with the first data information and the second data information, so as to more realistically restore the unique voice characteristics and the speaker's posture when speaking displayed in the meeting, and combine the speaker's voice and speaker behavior to restore and generate the video data corresponding to the video conference, so as to quickly and accurately convey the video data corresponding to the video conference to the receiving end, so that the receiving end can more accurately understand the speaker's emotions, tone and intentions, so as to improve the conference experience of the receiving end users while solving the problems faced by data transmission in weak network environments.
[0101] It can be understood that the above-mentioned device embodiment corresponds to the method embodiment of the present invention, which can implement any of the above-mentioned method embodiments of the present invention to provide a method for optimizing video conferencing quality in a weak network scenario.
[0102] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Furthermore, in the drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which may be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement the present invention without inventive effort.
[0103] Based on the above-mentioned embodiment of the method for optimizing video conferencing quality in a weak network scenario, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for optimizing video conferencing quality in a weak network scenario of any embodiment of the present invention is implemented.
[0104] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.
[0105] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0106] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.
[0107] Based on the above-mentioned method embodiments, another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the video conferencing quality optimization method in a weak network scenario described in any one of the above-mentioned method embodiments of the present invention.
[0108] Wherein, the module / unit integrated in the device / terminal equipment, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0109] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for optimizing video conferencing quality in a weak network scenario, characterized in that: include: During a video conference, first data information corresponding to a speaker's human behavior characteristics and the speaker's voice data are obtained in real time, and whether the current system is in a weak network scenario is determined in real time; wherein the first data information includes speaker identity information and speaker behavior characteristic information; If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information and the second data information, and combines the speaker voice and the speaker behavior to restore and generate the audio and video data corresponding to the video conference.
2. The method for optimizing video conferencing quality in a weak network scenario according to claim 1, wherein: If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information, and the second data information, specifically: If the current system is in a weak network scenario, converting the voice data into corresponding second data information, and transmitting the first data information and the second data information to a receiving end, so that the receiving end retrieves the speaker voiceprint information corresponding to the speaker identity information from a database based on the speaker identity information; The receiving end generates corresponding speaker voice according to the speaker voiceprint information and the second data information by using natural language processing technology and a speech synthesis engine.
3. The method for optimizing video conferencing quality in a weak network scenario according to claim 1, wherein: If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information, and the second data information, specifically: If the current system is in a weak network scenario, converting the voice data into corresponding second data information, and transmitting the first data information and the second data information to a receiving end, so that the receiving end calls a speaker portrait model corresponding to the speaker identity information from a database based on the speaker identity information; The receiving end generates a continuous action of the corresponding portrait model according to the portrait model and the speaker behavior feature information, and the speaker behavior is constituted based on the continuous action of the portrait model.
4. The method for optimizing video conferencing quality in a weak network scenario according to claim 3, wherein: The receiving end generates, according to the portrait model and the speaker behavior feature information, a continuous action of the corresponding portrait model, and constitutes the speaker behavior based on the continuous action of the portrait model, specifically: Based on the speaker's behavior feature information, obtaining a plurality of human behavior guiding words of the speaker and the order of all the human behavior guiding words; wherein the human behavior guiding words include expression guiding words and / or gesture guiding words; Based on a plurality of human behavior guiding words of the speaker, calling from a database the limb change sequences corresponding to the human behavior guiding words; wherein the limb change sequences are pre-generated using motion capture and animation generation technology and stored in the database; According to the order of all the human behavior guiding words, the limb change sequence corresponding to each of the human behavior guiding words is demonstrated on the portrait model in turn to form the continuous movement of the portrait model, and based on the continuous movement of the portrait model, the speaker behavior is constructed.
5. The method for optimizing video conferencing quality in a weak network scenario according to claim 1, wherein: If the current system is in a weak network scenario, the voice data is converted into corresponding second data information, and the first data information and the second data information are transmitted to the receiving end, specifically: If the current system is in a weak network scenario, corresponding voice data is obtained based on the voice signal, and then a pause and silence duration of the voice data is detected. Based on the detection result, the voice data is preprocessed to obtain a corresponding vector sequence; wherein the voice signal includes the voice data, and the vector sequence includes at least one voice vector; By calling a speech recognition model, speech recognition is performed on the speech vectors in the vector sequence, and second data information corresponding to the speech data is generated by combining the detection result and the speech recognition result.
6. The method for optimizing video conferencing quality in a weak network scenario according to claim 1, wherein: In the process of the video conference, the first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data are obtained in real time, specifically: During the video conference, human behavior characteristics corresponding to the speaker and the speaker's voice data are obtained in real time; wherein the human behavior characteristics include facial features and / or body movements; Comparing the facial features with pre-stored user data to obtain corresponding speaker identity information; Combining the facial features and the voice data, the speaker's emotional features are extracted in real time to obtain corresponding speaker emotional information, and key frame extraction and pattern matching are performed on the body movements to obtain corresponding special gesture signals. Then, the speaker's emotional information and the special gesture signals are combined to constitute the first data information corresponding to the speaker's human behavior characteristics.
7. A method for optimizing video conferencing quality in a weak network scenario according to any one of claims 1 to 6, characterized in that: The real-time determination of whether the current system is in a weak network scenario is specifically as follows: During the video conference, network indicator values corresponding to several system network indicators of the current system are obtained in real time, and each of the network indicator values is compared with a preset limit range of the corresponding system network indicator; If any of the network indicator values does not meet the preset limit range of the corresponding system network indicator, it is determined that the current system is in a weak network scenario; If all the network indicator values are within the preset limit range of the corresponding system network indicator, it is determined that the current system is not in a weak network scenario.
8. A video conferencing quality optimization system in a weak network scenario, characterized by: include: An acquisition and judgment module is used to obtain, in real time during the video conference, first data information corresponding to the speaker's human behavior characteristics and the speaker's voice data, and to determine in real time whether the current system is in a weak network scenario; wherein the first data information includes the speaker's identity information and the speaker's behavior characteristic information; A video restoration module is used to convert the voice data into corresponding second data information if the current system is in a weak network scenario, and transmit the first data information and the second data information to the receiving end, so that the receiving end calls the speaker voiceprint information corresponding to the speaker identity information, and then generates the corresponding speaker voice and speaker behavior based on the speaker voiceprint information, the first data information and the second data information, and combines the speaker voice and the speaker behavior to restore and generate the audio and video data corresponding to the video conference.
9. The video conferencing quality optimization system in a weak network scenario according to claim 8, characterized in that: The video restoration module specifically includes: an information transmission unit, configured to, if the current system is in a weak network scenario, convert the voice data into corresponding second data information, and transmit the first data information and the second data information to a receiving end, so that the receiving end retrieves the speaker voiceprint information and portrait model corresponding to the speaker identity information from a database based on the speaker identity information; The video restoration unit is used to generate the corresponding speaker voice based on the speaker voiceprint information and the second data information through the receiving end using natural language processing technology and a speech synthesis engine; generate the corresponding continuous movement of the portrait model based on the portrait model and the speaker behavior feature information through the receiving end, and constitute the speaker behavior based on the continuous movement of the portrait model; and restore and generate the audio and video data corresponding to the video conference in combination with the speaker voice and the speaker behavior.
10. A terminal, characterized in that: It includes a processor, a memory and a computer program stored in the memory; wherein, the computer program can be executed by the processor to implement a method for optimizing video conferencing quality in a weak network scenario as described in any one of claims 1 to 7.